Adaptive decision-making method based on cooperative interference of multiple unmanned aerial vehicles

By employing an adaptive decision-making method based on multi-agent reinforcement learning, the problem of interference resource conflict in multi-UAV cooperative interference was solved. This enabled UAVs to achieve cooperative optimization in complex electromagnetic environments, improving interference effectiveness and survivability while reducing energy consumption and risks.

CN121995326APending Publication Date: 2026-05-08CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING UNIV OF POSTS & TELECOMM
Filing Date
2026-01-16
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In multi-UAV cooperative interference scenarios, the lack of an effective cooperative decision-making mechanism leads to interference resource conflicts, redundancy superposition, and increased energy consumption. Furthermore, existing methods are not adaptable enough to dynamic environments and cannot effectively avoid interference resource conflicts between multiple UAVs.

Method used

An adaptive decision-making method based on multi-agent reinforcement learning is adopted to construct a multi-UAV cooperative interference task decision model. Through a centralized training-distributed execution (CTDE) multi-agent reinforcement learning framework, the UAVs achieve collaborative optimization in interference target selection, interference mode, timing and power allocation. Each UAV makes independent decisions based on its own observation information, avoiding redundant interference and resource conflicts.

Benefits of technology

Effectively avoid interference resource conflicts between multiple UAVs in complex electromagnetic environments, improve the overall interference effectiveness and swarm survivability of the system, reduce energy consumption and the risk of enemy radar lock-on, and enhance system-level interference effectiveness and UAV swarm survivability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121995326A_ABST
    Figure CN121995326A_ABST
Patent Text Reader

Abstract

The invention provides a self-adaptive decision-making method based on multi-unmanned aerial vehicle cooperative interference. The method comprises the following steps: constructing a multi-unmanned aerial vehicle cooperative interference task decision-making model in an electronic countermeasure scene containing a radar and multiple unmanned aerial vehicles; based on the multi-unmanned aerial vehicle cooperative interference task decision model, performing parametric modeling on the interference behavior of the interference unmanned aerial vehicle to generate an interference action space; training the multi-unmanned aerial vehicle cooperative interference task decision model based on the interference action space to obtain a cooperative interference decision model; a cooperative interference decision model is deployed on each interference unmanned aerial vehicle, each interference unmanned aerial vehicle extracts radar signal features and constructs a local observation state only based on radar radiation signals intercepted by an airborne electronic support measure system of the interference unmanned aerial vehicle, and an interference decision is independently output under the condition of no centralized control and no explicit communication, so that a cooperative interference task is completed; according to the method, through a global value evaluation mechanism in a centralized training stage, multiple unmanned aerial vehicles can form implicit cooperation in interference target selection, interference modes and interference power control levels in an execution stage, and redundant interference and interference resource conflicts of the multiple unmanned aerial vehicles on the same target are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the cross-technical field of intelligent decision-making, electromagnetic spectrum countermeasures, and cooperative control of unmanned aerial vehicle (UAV) swarms, specifically involving an adaptive decision-making method based on cooperative interference from multiple UAVs. Background Technology

[0002] With technological advancements, the battlefield of military confrontation has expanded beyond land, sea, and air, making the electromagnetic spectrum a hotspot in current great power rivalry. The increasing intelligence and informatization of military weapons mean that the outcome of electronic warfare will directly impact the course of war, and gaining an advantage in electronic warfare is crucial for ultimate victory. Electronic warfare refers to using electromagnetic energy and other means to simultaneously protect one's own equipment while hindering the normal operation of enemy electronic equipment; it is a primary form of current information warfare. Radar and jamming forces, as opposing sides in this confrontation, are constantly advancing their technologies in the process.

[0003] However, in multi-UAV cooperative jamming scenarios, without an effective collaborative decision-making mechanism, conflicts can easily arise among multiple jamming UAVs regarding target selection, jamming methods and timing, and power allocation. For example, multiple UAVs simultaneously launching high-power jamming against the same radar often leads to redundant and superimposed jamming resources, resulting in limited improvement in the overall jamming effectiveness of the system. Instead, it significantly increases energy consumption and the risk of detection and locking by enemy radar. At the same time, some high-threat radars may not be effectively suppressed due to unreasonable allocation of jamming resources, thereby reducing the overall survivability of the cluster. These phenomena are commonly referred to as "mutual interference" or "jamming resource conflict" among multiple UAVs in engineering practice.

[0004] In recent years, reinforcement learning and multi-agent reinforcement learning methods have been introduced into the field of electronic warfare decision-making. However, existing research mostly focuses on single jammers or simple cooperative scenarios, and most have failed to systematically address the problem of interference resource conflicts among multiple UAVs at the decision-making level. Some methods rely on preset rules or centralized control, which lacks adaptability in dynamic environments; others, while incorporating multi-agent learning, lack global constraints on cooperative behavior, potentially leading to redundant interference and resource waste. Therefore, there is an urgent need for a cooperative decision-making method that can proactively avoid ineffective interference between multiple UAVs in dynamic electromagnetic environments through learning mechanisms. Summary of the Invention

[0005] To address the problems existing in the prior art, this invention proposes an adaptive decision-making method based on multi-UAV cooperative jamming. This method includes: constructing a multi-UAV cooperative jamming task decision model for an electronic warfare scenario involving radar and multiple UAVs; parameterizing the jamming behavior of the jamming UAVs based on the multi-UAV cooperative jamming task decision model to generate a jamming action space; training the multi-UAV cooperative jamming task decision model based on the jamming action space to obtain a cooperative jamming decision model; and deploying the cooperative jamming decision model onto each jamming UAV. Each jamming UAV, based solely on the radar radiation signals intercepted by its own airborne electronic support measures system, extracts radar signal features and constructs a local observation state, independently outputting jamming decisions under conditions of no centralized control and no explicit communication, thereby completing the cooperative jamming task.

[0006] The beneficial effects of this invention are:

[0007] This invention addresses the dynamic game problem between multiple jamming drones and multiple radar systems. By constructing a multi-agent reinforcement learning decision-making framework based on centralized training-distributed execution (CTDE), it achieves collaborative optimization among multiple drones in areas such as target selection, jamming methods and timing, and power allocation without relying on real-time centralized control and explicit communication. This effectively avoids jamming resource conflicts among multiple drones in complex electromagnetic environments, improving the overall jamming effectiveness and swarm survivability of the system. Through a global value assessment mechanism in the centralized training phase, this invention enables implicit collaboration among multiple drones in the execution phase at the levels of target selection, jamming methods, and jamming power control, avoiding redundant jamming and jamming resource conflicts among multiple drones targeting the same target. It achieves collaborative jamming decision-making among multiple drones without relying on real-time centralized control and explicit communication, reducing system complexity and improving engineering feasibility. While reducing overall energy consumption and the risk of enemy radar lock-on, it significantly improves system-level jamming effectiveness and drone swarm survivability. Attached Figure Description

[0008] Figure 1 This is a schematic diagram of the scenario of the present invention;

[0009] Figure 2 This is the network architecture of the MA-CJD algorithm of the present invention;

[0010] Figure 3 This is a flowchart of the algorithm of the present invention. Detailed Implementation

[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0012] An adaptive decision-making method based on multi-UAV cooperative interference, such as Figure 2 As shown, the method includes: constructing a multi-UAV cooperative jamming task decision model in an electronic countermeasures scenario involving radar and multiple UAVs; parameterizing the jamming behavior of the jamming UAVs based on the multi-UAV cooperative jamming task decision model to generate a jamming action space; training the multi-UAV cooperative jamming task decision model based on the jamming action space to obtain a cooperative jamming decision model; deploying the cooperative jamming decision model to each jamming UAV, where each jamming UAV extracts radar signal features and constructs a local observation state based solely on the radar radiation signals intercepted by its own airborne electronic support measures system, and independently outputs jamming decisions under conditions of no centralized control and no explicit communication, thereby completing the cooperative jamming task.

[0013] To address the shortcomings of existing technologies, this invention proposes an adaptive optimization method for multi-UAV cooperative jamming decision-making based on multi-agent reinforcement learning (MA-CJD). This invention uses Markov game models as its theoretical foundation, modeling the multi-UAV cooperative jamming task as a multi-agent cooperation problem under partially observable conditions. It innovatively integrates the QMix value decomposition framework with the MP-DQN (Multi-Pass Deep Q-Network) parameterized action processing structure. By introducing global information and global rewards during the training phase, it uniformly optimizes the joint behavior of multiple UAVs, enabling the UAV swarm to form complementary rather than superimposed cooperative behaviors during the execution phase without real-time centralized control. This effectively avoids mutual interference between multiple UAVs at the decision-making level.

[0014] To achieve efficient coordinated jamming of multiple UAVs in complex electromagnetic environments, such as Figure 1 As shown, this invention constructs an electromagnetic countermeasures model for multiple radars and multiple jamming UAVs, and models the multi-UAV cooperative jamming decision-making problem as a partially observable Markov game. In this model, each UAV acts as an independent intelligent agent, which can only obtain local observation information through the onboard electronic support measures (ESM) system, and cannot directly know the internal operating status of the radar or the real-time decisions of other UAVs.

[0015] To address the potential interference and resource conflicts that may arise during multi-UAV collaboration, this invention introduces a centralized training-distributed execution (CTDE) multi-agent reinforcement learning framework. During the centralized training phase, the system utilizes global state and global rewards to uniformly evaluate the joint interference behavior of multiple UAVs, suppressing any individual behaviors that lead to a decrease in overall system interference effectiveness or resource waste during training. During the distributed execution phase, each UAV makes independent decisions based solely on its own observation information and the pre-trained policy model, forming a statistically consistent interference strategy without explicit communication or centralized scheduling. This naturally avoids ineffective interference and resource conflicts among multiple UAVs during the execution phase.

[0016] In this embodiment, to address the potential interference resource conflicts that may arise during multi-UAV collaboration, the present invention introduces a centralized training-distributed execution (CTDE) multi-agent reinforcement learning framework. During the centralized training phase, the system utilizes global state and global rewards to uniformly evaluate the joint interference behavior of multiple UAVs, suppressing any individual behavior that leads to a decrease in overall system interference effectiveness or resource waste during training. During the distributed execution phase, each UAV makes independent decisions based solely on its own observation information and the pre-trained policy model, forming a statistically consistent interference strategy without explicit communication or centralized scheduling. This naturally avoids ineffective interference and resource conflicts between multiple UAVs during the execution phase.

[0017]

[0018] in, For a collection of intelligent agents, there are individual jammers performing the jamming task; It represents the global state space, which includes the operating status (search, confirm, track) of all radars, their positions, the mainboard beam direction, the threat level, and the position information of all UAVs; Indicates the first The motion space of each drone is designed using parametric motion, as detailed below; Represents the state transition function. Describe the dynamic evolution of the environment under joint actions; For the reward function, The global reward provided by the system is used to guide the agent to learn cooperative strategies; It is a discount factor. .

[0019] Intelligent agents actually use observable states. Its characteristics are composed of signals measured by ESM of UAV:

[0020]

[0021] in This indicates an estimated approximate radar location; / / (Frequency / Pulse Width / PRI) indicates the frequency used to determine changes in radar mode; received power This indicates a significant change occurring during the main lobe scan; main lobe probability. The signal amplitude characteristics are used to estimate the signal; the historical sequence represents the time patterns of the input GRU capturing scan periodicity, irradiation stability, etc.

[0022] This invention does not assume that the UAV can see the radar status, but instead learns radar behavior implicitly through a network using time-series features: Search phase: Regular oscillations, power exhibiting scanning fluctuations; Confirmation phase: It briefly pauses in a localized area, and the fluctuations weaken; tracking phase: The behavior is basically stable, with power consistently and slightly higher than normal. Because it uses a GRU structure to encode historical observations, the agent can automatically infer these behaviors without needing the true internal state values.

[0023] The above modeling method ensures that the UAV does not rely on the actual state of the radar and the real-time actions of other UAVs when making decisions, so that the cooperative behavior is implicitly realized by the policy parameters learned during the training phase.

[0024] In this embodiment, to construct a high-fidelity simulation environment, the present invention establishes a detailed functional level model of radar and jammer.

[0025] 1) Radar main lobe scanning model:

[0026] In search mode, the main lobe beam direction of the radar antenna's reverse pattern changes periodically within a certain range. Main lobe beam direction at time The model is as follows:

[0027]

[0028] in, For radar main lobe width, For the scanning range, For scan time, For the number of scans, This is the pulse accumulation time.

[0029] 2) Target detection model:

[0030] The radar model calculates the detection probability based on the signal-to-noise ratio (SNR), using the Albertsheim approximation formula:

[0031]

[0032]

[0033]

[0034]

[0035] in, SNR is the false alarm probability. For a real target, SNR is defined as the ratio of echo power to the sum of noise power and suppression interference power. For a false target, SNR corresponds to the ratio of deception interference power to noise power.

[0036] 3) Interference signal reception and processing model:

[0037] The jammer employs two jamming methods: suppression jamming and deception jamming. Suppression jamming generates an interference background at the radar reception point, masking the echo signal and making it difficult for the radar to detect the real target. Deception jamming uses signals with characteristics similar to those of the radar to generate false target echoes, thereby misleading the radar. The power of the jamming signal received at the radar is defined as follows:

[0038]

[0039] in, This refers to the jammer's transmission power. This refers to the transmit antenna gain of the jammer. This represents the radar's receiving antenna gain in the direction of the jammer. For radar wavelength, The distance between the jammer and the radar. For the overall transmission loss of the jammer, For the overall receiving loss of the radar, For atmospheric loss, The instantaneous bandwidth of the radar receiver. This refers to the bandwidth of the interference signal. The radar has anti-jamming capabilities and employs an anti-jamming factor. Attenuate the received interference power . The value range is related to whether the interference signal comes from the main lobe or the side lobe, reflecting the type of radar and the level of anti-interference.

[0040] 4) Deception and interference identification model:

[0041] Furthermore, this radar can identify false targets generated by jamming based on the characteristics of the jamming signal. When the jitter frequency (JNR) is high enough, the signal characteristics become more pronounced. This leads to higher correlations between false target signals and lower correlations between false targets and real targets, thereby increasing the probability of identifying false targets. Based on this, the identification probability... and The relationship is modeled as a sigmoid function:

[0042]

[0043] in, The recognition probability is equivalent to gradient, This represents the JNR value with a recognition probability of 0.5.

[0044] 5) Calculation of signal-to-noise ratio under interference:

[0045] Undetected deception jamming signals will create false targets within the radar's main beam coverage area. The signal-to-noise ratio of the false targets is calculated as follows:

[0046]

[0047] The signal-to-noise ratio of the actual target is affected by jamming. Let the total power of the jamming signal received by the radar be... The radar's pulse compression accumulation gain is Under the influence of suppression interference, the signal-to-noise ratio of the real target within the main beam is as follows:

[0048]

[0049] in, and These represent the received suppression and deception interference power, respectively. Accumulate gain for radar pulse compression.

[0050] In this embodiment, during the cooperative jamming decision-making process, each jammer must determine its jamming target, jamming type, and power level at each step. This invention designs the jamming strategy as a parameterized action space consisting of discrete and continuous actions.

[0051] The jamming target and type are represented as discrete actions, while the jamming power level is represented as a continuous action. The jammer's actions are represented as follows:

[0052]

[0053] In the formula, Discrete variables represent the interference target and the type of interference. When When this occurs, it indicates that the jammer will not perform jamming actions. If so, it constitutes deception and interference; if This constitutes suppression interference. The target ID for interference is determined by... Sure, maximum value ( (Number of radar systems). Represents the continuous value of the jammer's power level, the normalized jamming power level, let and These represent the maximum and minimum transmit power of the jammer, respectively. The actual jamming power is then... With interference power level The relationship is defined as follows:

[0054]

[0055] By introducing continuous power parameters, multiple UAVs can automatically generate a division of labor between main and auxiliary interference during the execution phase, avoiding the waste of resources caused by multiple UAVs simultaneously interfering with the same target at high power.

[0056] In this embodiment, the multi-objective reward function design includes: during the collaborative interference decision-making process, the reward function not only reflects the immediate effect of a single interference action but also embodies the overall long-term benefits of the system. This invention employs a unified global reward function to evaluate the joint behavior of multiple UAVs. When multiple UAVs perform redundant interference on the same radar without significantly improving the interference effect, the global reward obtained by the system will decrease due to increased energy consumption and increased radar lock-on risk. In this way, multiple UAVs gradually learn to avoid ineffective superimposed interference on the same target during training, achieving a reasonable allocation of interference resources. This invention designs the reward function into three components: tracking penalty... Resource consumption penalty Probability of successful interference Therefore, the reward function is calculated as follows:

[0057]

[0058] Indicates radar lock-on penalty, As a penalty for resource consumption, A reward is given based on the probability of successful jamming. When a defense unit is locked and tracked by radar, a penalty is imposed based on the radar's threat level. This penalty incentivizes the agent to optimize jamming of high-threat radars and minimize the chance of the defense unit being locked on.

[0059]

[0060] in This is the penalty coefficient (usually a negative value). Indicates radar The threat level is determined based on prior information such as radar type, power, and anti-jamming capability. However, in actual systems, whether a radar is locked can be obtained through its own radar alarm receiver or collaborative sensing information.

[0061] This represents a resource consumption penalty. It is linearly related to transmission power, encouraging energy conservation. The maximum and minimum values ​​are respectively and , The calculation formula is as follows:

[0062]

[0063] In this invention, Set to -0.1, Set to -0.01. This represents the reward for the probability of successful interference. To alleviate sparse rewards, this reward provides dense feedback that is directly related to the effectiveness of the interference.

[0064] In multi-UAV cooperative jamming scenarios, the jamming effect typically exhibits significant nonlinear characteristics in relation to jamming power, timing, and the cooperative relationship among multiple UAVs. Simply using a linear superposition of several observed events is insufficient to accurately characterize the true contribution of a single UAV to the overall jamming mission. Therefore, this invention proposes a jamming effect model based on joint modeling of jamming effectiveness and radar behavior uncertainty. The reward calculation method is used to more effectively distinguish between effective interference and redundant interference during the intensive training phase, guiding multiple UAVs to form complementary and cooperative interference strategies.

[0065] For the A drone at all times The jamming effect function of the jamming action against the target radar is defined as follows:

[0066]

[0067] in This represents the change in the signal-to-noise ratio of the radar received signal under interference compared to the interference-free condition; This indicates the change in radar dwell time or tracking behavior in a specific direction relative to the historical average. , These are the weighting coefficients.

[0068] The utility function described above has the characteristics of monotonically increasing and diminishing marginal returns. It is used to reflect that when multiple UAVs simultaneously interfere with the same radar, the interference effect does not increase linearly with power or the number of interferences, thereby suppressing redundant interference behavior in terms of reward layer.

[0069] Based on historical observation sequences acquired by the ESM system without relying on the internal state of the radar. The intelligent agent can form a probability distribution estimate of the radar's current behavior model. Its certainty is measured by information entropy:

[0070]

[0071] Define the change in radar behavior uncertainty as:

[0072]

[0073] When interference causes frequent switching of radar operating modes, unstable beam direction, or unpredictable behavior characteristics, this uncertainty increases, reflecting the effectiveness of interference in the radar's decision-making process.

[0074] Taking into account both the interference effect and the uncertainty of radar behavior, the first [missing information] in this invention Rewards for interference effects of drones Defined as:

[0075]

[0076] in This is the uncertainty reward weighting coefficient. This jamming effect reward depends only on the characteristics of the ESM observable signal and its temporal changes, without requiring information about the radar's actual internal operating state. It can effectively distinguish between jamming behaviors that make a substantial contribution to the overall jamming mission and redundant jamming behaviors that have limited improvement on the jamming effect during the intensive training phase.

[0077] This invention employs a centralized training-distributed execution multi-agent reinforcement learning training approach. The centralized training phase occurs before algorithm deployment and is typically completed in a simulation environment or ground computing platform. During this phase, the system has access to global state information and global rewards, and performs unified optimization of the joint behavior of multiple UAVs through the QMix value decomposition mechanism. After training is completed, the policy model parameters of each UAV are fixed and deployed to the UAV platform.

[0078] In the actual execution phase, each UAV no longer undergoes centralized training or online learning. Instead, it independently makes interference decisions based solely on its own ESM observation information and by invoking the pre-trained policy model. Since the cooperative behavior has been internalized into the policy parameters through global optimization during the centralized training phase, even without real-time information interaction between UAVs in the distributed execution phase, a coordinated and consistent interference behavior can still be formed as a whole.

[0079] 1) Training framework

[0080] In multi-agent systems, the policy updates of each interfering agent influence each other, potentially leading to environmental non-stationarity and reward misdirection. To overcome this problem, this invention employs the QMix algorithm with a centralized training and distributed execution (CTDE) model. Each agent... Possesses an independent action value network This is used to predict the value of its own actions; at the same time, a global mixing network is set up during the training phase:

[0081]

[0082] The network uses the local action Q-values ​​and global states of all agents. The input is the global action value. Network parameters are constrained to be non-negative to ensure the monotonicity of both global and individual values, thus achieving decomposable consistency between individual and global optima. Through this structure, the system can automatically allocate the overall task value to each interfering machine during the training phase, avoiding individual reward bias or "laziness," and improving training stability and global synergy.

[0083] 2) Network Structure

[0084] MP-DQN Network: To address the joint action space containing discrete actions and continuous parameters, this invention employs the MP-DQN (Multi-Pass Deep Q-Network) structure. Its core idea is: discrete actions...

[0085] Corresponding interference targets and interference types; continuous parameters This indicates the normalized interference power level.

[0086] MP-DQN consists of two parts:

[0087] 1. Actor Network: Based on the state vector Generate power parameters for each discrete action. ;

[0088] 2. Q-Network: By combining the state, discrete actions, and corresponding power parameters as input, the value of each action is calculated.

[0089]

[0090] The power parameters output by the Actor are concatenated with the state and then input into a multi-channel Q-network. The hidden layer of the Q-network uses GRU units to fuse historical state and action information, achieving temporal decision association. Action selection is based on an ε-greedy strategy, and its output action value... Global action value is formed through aggregation via hybrid networks:

[0091]

[0092] QMix Hybrid Network: The hybrid network consists of a hyper-network and is used to decompose and aggregate the global Q-value. Its output satisfies the monotonicity constraint.

[0093]

[0094] Ensure that increasing individual value does not decrease the overall system value. Hypernetwork weight parameters are determined through the global state. Dynamic generation enables the system to adaptively adjust the collaborative weights among agents based on changes in the situation.

[0095] In this embodiment, the overall architecture of the UAV cooperative jamming decision system mainly includes the following components: a scenario modeling module for establishing a multi-radar-multi-jammer electromagnetic countermeasures simulation environment; a state perception module for extracting dynamic feature information of the radar and jammer; a strategy decision module for outputting jamming targets, jamming methods, and power control commands; a network training module for executing a multi-agent reinforcement learning algorithm with centralized training and distributed execution (CTDE); and a simulation verification module for verifying the effectiveness of the jamming decision.

[0096] The system achieves cooperation and adaptive optimization among jammers through the above mode, thereby maximizing the jamming effect under energy constraints.

[0097] In this embodiment, scene modeling and state construction are shown in Table 1.

[0098] Table 1 Key Parameters for Different Radar Types

[0099]

[0100] The model construction includes an electronic countermeasures scenario with 4 UAVs and 4 radars. The parameters of each radar are shown in Table 1, with different transmit power, main lobe width, and anti-jamming factor. The jamming UAVs are initially distributed around the target area, 50km away from the center point, with azimuth angles of 45°, -45°, 135°, and -135°.

[0101] The "4-on-4" scenario adopted in this invention is the minimum typical configuration for algorithm development and functional verification. The MA-CJD algorithm architecture proposed in this invention has good scalability. Its core centralized training-distributed execution (CTDE) framework, parameterized action space design, and credit allocation mechanism all support expanding the jamming unit scale to a swarm of dozens or more UAVs to cope with large-scale, high-density electromagnetic countermeasures scenarios involving more radar nodes. In practical engineering applications, collaborative decision-making by large-scale swarms can be achieved by adjusting the capacity of the mixing network and optimizing the distributed communication protocol.

[0102] Radar detection states include three phases: "search—confirm—track," such as... Figure 3 As shown.

[0103] In this embodiment, the radar model calculates the detection probability based on the signal-to-noise ratio (SNR), using the Albersheim approximation formula:

[0104]

[0105] in , SNR is the false alarm probability. For a real target, SNR is defined as the ratio of echo power to the sum of noise power and suppression interference power. For a false target, SNR corresponds to the ratio of deception interference power to noise power.

[0106] Example 3: Interference model and signal power calculation.

[0107] The jammer employs two jamming methods: active suppression jamming and active deception jamming. The power of the jamming signal received by the radar is defined as:

[0108]

[0109] And introduce anti-interference factors This indicates that the effective interference power after radar anti-jamming processing is:

[0110]

[0111] The probability of identifying deception interference and The relationship between the interference-to-noise ratio and the noise ratio is shown below:

[0112]

[0113] The formula for calculating the signal-to-noise ratio of a real target under suppression interference is as follows:

[0114]

[0115] The formula for calculating the signal-to-noise ratio of a false target under deception interference is as follows:

[0116]

[0117] Example 4: Design of Parametric Action Space and Multi-Objective Reward Function

[0118] This invention designs the action of each jammer as a parameterized action space:

[0119]

[0120] in For discrete actions, it represents the target of interference and the type of interference (deception / suppression). This represents the continuous power ratio. The actual power is:

[0121]

[0122] The reward function is designed as follows:

[0123]

[0124] in Penalty for being locked on by radar; This is a penalty for resource consumption and is linearly related to power. To interfere with the success probability reward.

[0125] Example 5: Reinforcement Learning Training Structure and Process

[0126] The training in this invention adopts a centralized training-distributed execution architecture under the QMix framework:

[0127] 1) Each agent has its own independent Actor network and Q network;

[0128] 2) Mixing Networks are used to aggregate the Q-values ​​of all agents and output the global Q-value. ;

[0129] 3) Employ the Double DQN mechanism to prevent overestimation of the Q value;

[0130] 4) Network parameters are updated using the Temporal Difference (TD) method, with the loss function being:

[0131]

[0132] in:

[0133]

[0134] The Actor network is a three-layer fully connected structure (128-128-1), with GRU units embedded in the Q network to remember state sequences. The Mixing network consists of two supernetwork layers, with weights adjusted by the global state to achieve adaptive credit allocation.

[0135] Example 6: Simulation and Performance Verification

[0136] Validation was performed on a C++ and Python co-simulation platform. The training environment parameters are as follows: learning rate... , Discount factor Exploration rate The value decreases from 0.95 to 0.05; each round of interaction consists of 16 episodes, and the experience pool batch size B=32.

[0137] Training 10 5After one step, the algorithm converges. To comprehensively evaluate the performance of the MA-CJD algorithm, we compare it with three representative baseline algorithms:

[0138] 1) Rule-based Policy: A heuristic strategy based on expert knowledge. Its core rules are: each UAV prioritizes jamming the radar with the highest threat level and located in its main lobe direction; if more than one radar meets the conditions, they are evenly distributed; the jamming mode is dynamically selected based on distance (suppression for close range, deception for long range); the transmit power is set to a fixed empirical value (50% of maximum power). This strategy represents a reasonable jamming logic based on fixed rules, requiring no online learning.

[0139] 2) PER-DDQN: An advanced single-agent deep reinforcement learning algorithm that employs Priority Experience Replay (PER) and Dual Deep Q-Network (DDQN) techniques. In this experiment, it treats multiple drones as a "joint agent," with the action space being the Cartesian product of all drone actions, facing a severe curse of dimensionality problem.

[0140] 3) Standard QMix: A classic multi-agent reinforcement learning algorithm that employs a centralized training and distributed execution framework. In this experiment, its action space is a discrete power level (only low, medium, and high levels), which cannot achieve continuous and fine-grained power control.

[0141] Table 2 shows a comparison of the average performance of each algorithm under the same test scenario:

[0142] Table 2 Comparison of Algorithms

[0143]

[0144] The proposed MA-CJD algorithm significantly outperforms all baseline algorithms. Compared to rule-based methods, MA-CJD achieves approximately 46% performance improvement, demonstrating that intelligent learning strategies far surpass heuristic methods based on fixed rules. Compared to PER-DDQN, MA-CJD's significant advantage highlights the effectiveness of multi-agent methods in solving collaborative problems, avoiding the dimensionality curse of single-agent methods in many-to-many scenarios. Compared to the standard QMix, MA-CJD's performance leap is mainly attributed to its parameterized action space design (integrating MP-DQN), enabling agents to continuously and finely adjust interference power, rather than being limited to a few discrete power levels, thus achieving a better balance between interference effectiveness and resource efficiency.

[0145] Simulation results show that the method of the present invention can effectively avoid the situation where interference resources are concentrated on a single target in multi-UAV cooperative scenarios, verifying the effectiveness of the centralized training-distributed execution mechanism in suppressing invalid interference from multiple UAVs.

[0146] The invention can be extended to: real-time jamming decision-making in UAV swarm electronic warfare missions; power control and task allocation in distributed jamming networks; and other multi-agent warfare scenarios, such as the coordinated scheduling of unmanned surface vessels or ground jamming vehicles. This method supports flexible deployment under different frequency bands, power constraints, and radar types, and possesses good scalability and engineering feasibility.

[0147] In summary, a multi-agent electromagnetic countermeasures (EMIC) cooperative jamming multi-dimensional joint decision-making algorithm (MA-CJD) based on multi-agent reinforcement learning is proposed. A mathematical model is established to represent the cooperative jamming process, combining radar and jamming UAV models. A Markov game framework is introduced, where the jamming target, model, and power calculation level are considered as decision variables in a parameterized action space. The QMix algorithm is combined with the MP-DQN architecture to improve learning efficiency and decision performance. Compared with existing methods, the MA-CJD algorithm significantly improves jamming performance while reducing resource consumption. This method proposes a new framework for multi-dimensional jammer parameter optimization and verifies the effectiveness of combining QMix and MP-DQN in cooperative decision-making tasks. Future work will explore more detailed models and incorporate additional parameters to enhance practical applicability.

[0148] The above-described embodiments further illustrate the purpose, technical solution, and advantages of the present invention. It should be understood that the above-described embodiments are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made to the present invention within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An adaptive decision-making method based on multi-UAV cooperative interference, characterized in that, include: Construct a multi-UAV cooperative jamming mission decision model in an electronic warfare scenario that includes radar and multiple UAVs; The multi-UAV cooperative jamming task decision model is used to parameterize the jamming behavior of the jamming UAVs and generate a jamming action space. The multi-UAV cooperative jamming task decision model is trained based on the jamming action space to obtain a cooperative jamming decision model. The cooperative jamming decision model is deployed on each jamming UAV. Each jamming UAV extracts radar signal features and constructs a local observation state based solely on the radar radiation signals intercepted by its own airborne electronic support measures system. Under conditions of no centralized control and no explicit communication, it independently outputs jamming decisions, thereby completing the cooperative jamming task.

2. The adaptive decision-making method based on multi-UAV cooperative interference according to claim 1, characterized in that, The decision-making model for multi-UAV cooperative jamming missions includes: modeling each jamming UAV as an independent decision-making agent, and modeling the multi-UAV cooperative jamming mission as a partially observable Markov game model; setting the radar's operating modes, including search mode, tracking mode, and confirmation mode; passively intercepting radar radiation signals through the airborne electronic support measures system according to the operating modes, obtaining the radar signal's angle of arrival, operating frequency, pulse width, pulse repetition interval, and received signal power, and constructing the local observation state of the jamming UAVs.

3. The adaptive decision-making method based on multi-UAV cooperative interference according to claim 2, characterized in that, Search mode: Regular oscillations, power exhibiting scanning fluctuations; Confirmation mode: The fluctuations weakened after a brief local pause; tracking mode: The power output is basically stable, but remains consistently high.

4. The adaptive decision-making method based on multi-UAV cooperative interference according to claim 1, characterized in that, Parametric modeling of the jamming behavior of jamming drones includes: representing the jamming behavior of the jamming drone at any decision moment as a jamming action consisting of discrete actions and continuous parameters; and calculating the actual transmission power of the jamming drone based on the jamming action.

5. The adaptive decision-making method based on multi-UAV cooperative interference according to claim 4, characterized in that, The actual transmit power of the jamming drone is calculated as follows: ; in, and These represent the maximum and minimum transmit power of the jammer, respectively; This represents a continuous value of the jammer's power level, a normalized jamming power level.

6. The adaptive decision-making method based on multi-UAV cooperative interference according to claim 1, characterized in that, The centralized training of the multi-UAV cooperative jamming mission decision model includes: acquiring a training set, introducing globally observable environmental information during the training phase, and constructing a model that includes the actual operating state of the radar. Where 0 represents the search state, 1 represents the confirmation state, and 2 represents the tracking state; [This indicates] acquiring the radar position. Maintaining relative stability and radar threat level under confirmed or tracked conditions. Information and the locations of each jamming drone The global state of information; a system-level reward function is constructed based on the global state and the joint jamming actions of multiple UAVs; when jamming causes frequent switching of radar operating modes, unstable beam direction or unpredictable behavioral characteristics, this uncertainty increases, reflecting the effective jamming of the jamming radar decision-making process. The above reward function is used to guide the learning process of the multi-UAV cooperative jamming strategy.

7. The adaptive decision-making method based on multi-UAV cooperative interference according to claim 6, characterized in that, The system-level reward function is: ; ; ; ; ; in, To indicate the penalty for radar lock-on, As a penalty for resource consumption, To interfere with the success probability reward, To estimate the probability distribution of radar behavior under historical observations, The uncertainty entropy of the radar's behavior at time t, Let be the interference utility function. For uncertain reward weighting coefficients, The variable representing the uncertainty of radar behavior. , These are the weighting coefficients. This represents the change in the signal-to-noise ratio of the radar received signal under interference compared to the interference-free condition. This indicates the change in radar dwell time or tracking behavior in a specific direction relative to the historical average.

8. The adaptive decision-making method based on multi-UAV cooperative interference according to claim 6, characterized in that, Value assessment of multi-UAV joint interference behavior includes: constructing a joint value function based on the global state and the joint interference actions of multiple UAVs; the network uses the local actions of all agents. Values ​​and global state As input, the global action value is output; the joint action value function is constrained to satisfy the monotonicity condition for the individual action value functions of each interfering UAV; an interference utility function is constructed based on changes in radar signal-to-noise ratio, radar tracking performance, and radar behavior uncertainty to quantitatively evaluate the joint interference effect of multiple UAVs; when the marginal gain of the interference utility function with respect to interference power is lower than a preset threshold... When this happens, the corresponding interference behavior is determined to be redundant interference, and its weight in the joint action value assessment is reduced.

9. The adaptive decision-making method based on multi-UAV cooperative interference according to claim 1, characterized in that, The distributed execution phase includes: each jamming UAV passively constructs its local observation state based solely on radar radiation signals passively intercepted by its own onboard electronic support measures system; the UAVs then use the cooperative jamming decision-making model obtained during the centralized training phase, based on their local observation states... The interference actions are calculated independently; there is no explicit communication of state, action, or value information between the interference drones, and the interference actions of each interference drone jointly constitute the joint interference behavior of the system; there is no centralized training and parameter update, and the cooperative interference effect is obtained from the policy consistency obtained by implicit learning in the centralized training phase.