A Multi-Agent, Multi-System Cooperative Detection Method Based on Deep Reinforcement Learning for Radar Jamming Countermeasures
By constructing a multi-agent, multi-system collaborative detection method based on deep reinforcement learning and employing the HyPPO-MACAD hybrid attention proximal policy optimization algorithm, the problem of anti-interference policy generation and path planning in complex dynamic game adversarial scenarios for multi-agent collaborative detection is solved, achieving efficient regional coverage and detection efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TONGJI UNIV
- Filing Date
- 2026-01-30
- Publication Date
- 2026-06-02
AI Technical Summary
Existing multi-agent cooperative detection technologies struggle to generate efficient anti-interference strategies and plan cooperative paths in complex dynamic game-based adversarial scenarios, and fail to effectively cope with time-varying and irregular radar interference environments, resulting in detection redundancy and excessive computational burden.
A multi-agent, multi-system collaborative detection method based on deep reinforcement learning is constructed. The HyPPO-MACAD hybrid attention proximal policy optimization algorithm is adopted, which combines a time-varying anisotropic radar detection model and a distributed partially observable Markov decision process. A continuous-discrete hybrid branch and attention mechanism are introduced to optimize the policy network to reduce detection stickiness and path conflict.
It improves the detection efficiency and anti-interference capability of multi-agent systems in dynamic and highly adversarial environments, reduces detection redundancy and computational burden, and achieves efficient regional coverage.
Smart Images

Figure CN122131248A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of electronic information technology and artificial intelligence, and specifically relates to a collaborative detection method in complex electromagnetic countermeasures environments, particularly a multi-agent, multi-system collaborative detection method based on deep reinforcement learning under radar jamming countermeasures. Background Technology
[0002] Cooperative reconnaissance refers to the planning of flight paths to reconnoiter a designated area of interest, aiming to maximize certain performance characteristics under given constraints. It has been widely applied in disaster relief, agricultural plant protection, and battlefield reconnaissance. With the upgrading and iteration of software and hardware technologies and the development of multi-agent technology, reconnaissance scenarios are exhibiting a complex trend towards dynamic, adversarial, and swarm-like characteristics, posing new challenges to reconnaissance algorithms. In military scenarios, dynamic adversarial characteristics are particularly prominent. Therefore, researching cooperative reconnaissance algorithms with online intelligent decision-making and anti-interference capabilities for complex and highly adversarial environments has become one of the core issues.
[0003] Currently, technical design schemes for collaborative region detection and coverage in this field mainly fall into two categories: those based on traditional rule-based or heuristic algorithms and their variants, and those based on deep reinforcement learning. Patent CN115031736A proposes a multi-UAV collaborative region coverage trajectory planning method and system based on geometric region division. By dividing the total search capability of the rendezvous point into a dual proportional division of the search capability of each UAV, it achieves multi-level task allocation, balancing global collaboration with individual efficiency. Simultaneously, it introduces trapezoidal decomposition and dynamic programming methods to reorganize complex concave polygons into convex polygon sub-regions, significantly reducing redundant flights and turning times, and improving search efficiency. Patent CN119091287A proposes a collaborative detection method based on information graph updates and local game theory for distributed multi-autonomous underwater vehicle underwater region detection. It introduces optimal dynamic response to generate local target points and achieves efficient group collaboration through anti-cluster control, avoiding redundant detection. Patent CN114879742A proposes a dynamic coverage method for UAV swarms based on multi-agent deep reinforcement learning. It introduces a centralized action corrector during the training phase to maintain connectivity, and combines this with a distributed execution framework to achieve dynamic target coverage. This supports dynamic expansion of the number and type of UAVs and robustness to individual UAV failures. Patent CN119044899A proposes a multi-agent reinforcement learning-based radar anti-jamming method based on jamming capability allocation. This method solves the problem of limited radar swarm communication during detection and effectively completes collaborative anti-jamming decision-making in dynamic game-based adversarial scenarios involving radar jammers.
[0004] However, the technical solution proposed in patent CN115031736A mainly targets offline region division in static, barrier-free environments, failing to effectively solve the real-time planning problem of multiple UAVs in highly dynamic adversarial environments. Furthermore, the task region hierarchical division in this patent, based on dynamic programming reorganization and cost matrix cyclic exchange, leads to high computational burden and limited real-time performance in large-scale multi-UAV scenarios. The solution proposed in patent CN119091287A has certain advantages in balancing computational overhead while optimizing group collaboration in a distributed manner, but its limitation lies in its region gridding and information map updates based on an ideal detection model, without considering the impact of complex obstacles or dynamic water flow on the detection model, making it unable to handle communication-constrained and dynamically changing game-theoretic detection tasks. Although the solution proposed in patent CN114879742A achieves specified coverage density for multiple targets through path planning, its constraint on communication connectivity increases computational complexity, potentially extending training time to some extent, and it does not explicitly consider moving obstacles or sudden environmental changes, making it unsuitable for large-scale, complex, dynamic game-theoretic scenarios. The solution proposed in patent CN119044899A completes the phased modeling of the dynamic game process of radar jammers. Based on this, the proposed multi-agent deep reinforcement learning algorithm effectively improves the radar's cooperative anti-jamming detection capability and autonomous decision-making capability. However, it only considers the jamming process of a single jammer on multiple radars and is mainly applicable to the static resource management and mode switching level. It fails to effectively address the complex decision-making requirements of multi-agent cooperative detection path planning and anti-jamming strategy generation in dynamic and complex environments.
[0005] Through research and analysis of existing technical solutions in environmental modeling, coupled decision learning, and online policy optimization for multi-agent cooperative game detection problems, the following technical bottlenecks were identified: Simple environmental constraints: Current multi-agent cooperative detection technologies mostly consider simple, unobstructed environments and ideal rule-based detection models, paying less attention to the time-varying and irregular characteristics of radar sensors in interference and countermeasure scenarios under game-based adversarial scenarios, thus resulting in a significant gap with practical applications.
[0006] Coupled Decision-Making and Online Learning Problems: Current deep reinforcement learning cooperative detection schemes typically consider single-type policies, either continuous or discrete. However, in adversarial game scenarios, where high real-time policy generation is required, how to achieve coupled decision-making for anti-interference policy generation and cooperative path planning in multi-agent cooperative detection problems, and how to study effective cooperative mechanisms to improve coverage efficiency of the task area and reduce detection redundancy, still require further design and optimization.
[0007] Information incompleteness in game theory: In complex and dynamic game environments, the environmental information acquired by a single agent is usually local and incomplete. How to enable effective collaboration among multiple agents within a team to gain a game advantage and make optimal decisions under such circumstances is one of the problems that current technologies struggle to solve. Summary of the Invention
[0008] The purpose of this invention is to provide a multi-agent, multi-system cooperative detection method based on deep reinforcement learning for radar jamming countermeasures. This method enables efficient cooperative detection of multiple agents under various anti-jamming mechanisms in highly contested battlefield environments with multiple radar jamming systems, thereby improving the agents' anti-jamming capabilities and detection efficiency. Specifically: This invention provides a multi-agent, multi-system cooperative detection method based on deep reinforcement learning for radar jamming countermeasures, characterized by the following steps: Step 1, constructing a dynamic adversarial game model of an aircraft cluster and a jamming aircraft cluster in a jamming scenario; Step 2, treating the jamming aircraft cluster as the environment, and modeling the interaction process between multiple aircraft and the environment as a distributed partially observable Markov decision process; Step 3, based on the distributed partially observable Markov decision process, using the HyPPO-MACAD multi-agent cooperative anti-jamming detection algorithm based on hybrid attention proximal policy optimization for training and decision-making to obtain a cooperative policy; the HyPPO-MACAD algorithm introduces a continuous-discrete hybrid branch in the policy network and an attention mechanism in the commentator network; Step 4, based on the trained cooperative policy, controlling the aircraft cluster to perform cooperative detection tasks.
[0009] Preferably, in step 1, the dynamic adversarial modeling of the aircraft swarm and the jamming aircraft swarm is performed using a tuple. The game problem is represented by, where For environmental maps, Gathering for players, and They are state space and action space, respectively. For the policy function, For the environment state transition function, For the sake of game gains.
[0010] Preferably, the game payoff Based on the calculation of the time-varying anisotropic radar detection model, the time-varying anisotropic radar detection model is used to calculate the point of the aircraft under multiple interference conditions. Detection range in the direction: In the formula, This represents the maximum detection range under interference-free conditions.
[0011] Preferably, in step 2, the interaction process between multiple aircraft and the environment is modeled as a series of tuples. The defined distributed partially observable Markov decision process, wherein, For aircraft clusters, For state space, It is the observation space. For the action space, Let be the state transition probability function. It is a discount factor. It is a reward function.
[0012] Preferably, the reward function A hierarchical incentive mechanism combining global and individual adaptive rewards is adopted, with individual rewards... ,in To effectively cover rewards, Adaptive penalty for repetitive regions. This is a reward for preventing interference.
[0013] Preferably, the formula for calculating the adaptive penalty for repeated regions is: In the formula, For grid Total number of visits.
[0014] Preferably, in step 3, the continuous-discrete hybrid branch of the policy network specifically refers to: the policy network It consists of two parts: a continuous branch and a discrete branch, both of which share two fully connected layers. The continuous branch is constructed by a fully connected layer and a Beta distribution layer. The discrete branch is constructed by a fully connected layer and a Category distribution layer.
[0015] Preferably, in step 3, the attention mechanism in the critic network specifically refers to: the critic network Includes a hybrid attention module for aggregating motion information from other aircraft and calculating the attention scores of other aircraft for the current aircraft. In the formula, , .
[0016] Preferably, in step 3, the HyPPO-MACAD algorithm employs a centralized training and distributed execution framework, utilizing gradient ascent to optimize the objective function through a proximal strategy. Update policy network parameters: .
[0017] Corresponding to the above method, the present invention also provides an electronic device for implementing the method: including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it can realize the above-mentioned multi-agent cooperative detection method based on deep reinforcement learning under radar jamming countermeasures.
[0018] Technical effect This invention models a novel game-theoretic adversarial model under radar jamming scenarios, combining it with a time-varying anisotropic radar detection model under jamming conditions to adapt to dynamic, highly adversarial area detection and target perception tasks.
[0019] This invention proposes a multi-agent cooperative anti-interference detection algorithm based on hybrid attention proximal policy optimization. It introduces a continuous-discrete hybrid branch in the policy network to adapt to the coupled decision-making task of anti-interference motion control, and introduces attention in the commentator network to extract global motion features, thereby reducing detection stickiness and path conflict among multi-agents. Furthermore, by designing a hierarchical adaptive reward to balance the global detection task and the stickiness penalty in the local range, it improves the efficiency of group cooperation and detection.
[0020] Through experimental simulation and performance analysis, the method of this invention can adapt to coupled decision-making tasks, perform intelligent anti-interference and collaborative detection, reduce detection adhesion under irregular detection models, and improve anti-interference capability and regional coverage efficiency. Attached Figure Description
[0021] Figure 1 This is a flowchart of a multi-agent, multi-system cooperative detection method based on deep reinforcement learning under radar jamming countermeasures in an embodiment of the present invention. Figure 2 This is a schematic diagram of a multi-agent, multi-system collaborative active game detection problem scenario under interference and adversarial conditions in an embodiment of the present invention. Figure 3 This is a schematic diagram of the detection range under different interference scenarios in the embodiments of the present invention; Figure 4 This is a structural diagram of the hybrid action space policy network (left) and the hybrid attention value network (right) in an embodiment of the present invention; Figure 5 This is a framework diagram of a multi-agent cooperative anti-interference detection algorithm with hybrid attention proximal strategy optimization in an embodiment of the present invention; Figure 6 This is a comparison chart of the algorithm training reward curves in an embodiment of the present invention; Figure 7 The following is a comparison of the collaborative detection performance of different algorithms under different agent scales in the embodiments of the present invention: region coverage (left) and average number of detections per covered grid (right). Detailed Implementation
[0022] To make the technical means, creative features, objectives and effects of this invention easy to understand, the following embodiments, in conjunction with the accompanying drawings, specifically illustrate a multi-agent, multi-system cooperative detection method based on deep reinforcement learning under radar jamming countermeasures.
[0023] Overall technical concept and modeling This section provides a multi-agent, multi-system cooperative detection method based on deep reinforcement learning for radar jamming countermeasures, to illustrate the overall technical process of the present invention.
[0024] This embodiment presents a multi-aircraft cooperative anti-jamming decision-making method based on improved deep reinforcement learning to maximize area detection performance in dynamic jamming environments. First, a dynamic adversarial game model of aircraft and jammers in a jamming scenario is constructed. A time-varying anisotropic radar model is modeled in detail based on the fundamental radar equations and the active jamming mechanism, serving as the game payoff to simulate area detection effectiveness under adversarial conditions. Then, the jammer is considered as part of the environment, and the multi-agent, multi-system cooperative game detection problem is formalized as a distributed, partially observable Markov decision process, defining each component. Finally, a multi-agent cooperative detection decision-making algorithm based on hybrid attention-based proximal policy optimization is proposed. A dual-branch approach is introduced into the policy network to adapt to the coupled decision problem, while an attention mechanism is added to the critic network to improve the rationality of global motion planning and reduce detection redundancy. Dual guidance from the behavior layer and reward layer promotes algorithm convergence and adaptability to dynamic and complex decisions. Details are as follows: Figure 1 This is a flowchart of a multi-agent, multi-system cooperative detection method based on deep reinforcement learning under radar jamming countermeasures in an embodiment of the present invention.
[0025] like Figure 1 As shown, the multi-agent, multi-system cooperative detection method based on deep reinforcement learning under radar jamming countermeasures in this embodiment includes the following steps: Step S1: Construct a dynamic adversarial game model of aircraft clusters and jamming aircraft clusters in the jamming scenario.
[0026] First, the following is the technical architecture of the multi-to-multi game adversarial model between radar aircraft and jammers.
[0027] Figure 2 This is a schematic diagram of a multi-agent, multi-system collaborative active game detection problem scenario under interference and adversarial conditions in an embodiment of the present invention.
[0028] In countermeasures-based area detection missions, the confrontation between radar-equipped aircraft swarms and jamming aircraft swarms is a dynamic, continuous, and highly uncertain game process, such as... Figure 2As shown. The detector aims to maximize its detection capability over the monitored area through pose control and effective anti-jamming strategies; while the jammer strives to minimize the radar's effective coverage area through distributed jamming strategies, thereby protecting its own information. Therefore, this section models the dynamic confrontation process between the radar and the jammer as a tuple. This represents a game theory problem. Among them... For environmental maps, Gathering for players, and They are state space and action space, respectively. The policy function is a mapping from the state space to the action space. For the environment state transition function, The outcome of the game is the direct basis for the actions of both parties.
[0029] Environment and Player Set: The quest area within a given environment ,exist A homogeneous aircraft carrying radar sensors and A fixed source of interference. The set of players participating in the game is... = Including the detection method Table and interference Two opposing camps.
[0030] State space: This refers to the task time. Discretization into a fixed and finite set of decision moments, i.e., At the moment of decision-making The joint state of the system Defined as the set of all player states: in, Let be the state vector for all aircraft. For each aircraft... Its state = Includes its location ,speed and heading angle . = Let be the state vector of all jammers. For a jammer... Its state = Includes its location .
[0031] Operational space and strategy: Each aircraft is equipped with radar. A kind of anti-interference system, using ={ , ,…, } indicates radar A collection of anti-interference mechanisms. For their strategy It is a mapping from state to action. .action Including motion control and anti-interference strategies Two parts, of which Satisfying constraints , ,in and Indicates the minimum and maximum acceleration of the aircraft. and Let represent the minimum and maximum values of the spacecraft. Therefore, the joint strategy of the detection parties is: Joint action for .use and = { , ,…, } indicates a jammer The total interference power and interference action set, among which This refers to the number of optional jamming systems. Jammer strategy It is a mapping from state to action. Considering the interference source in a fixed location, the actions of the jammer... Including power allocation vector and interference strategy selection vector .in The combined strategy of the interfering parties is as follows: Joint action for In the area detection mission described in this embodiment, it is assumed that the jammer's strategy is generated as follows: Given a total jamming power, the jammer allocates the jamming power based on the relative distance between the aircraft and the jammer: in, jammer With aircraft At any moment The distance. Furthermore, the jammer generates random jamming strategies for each aircraft based on random sampling.
[0032] State transition: The state transition of the system is described by both deterministic dynamics and disturbance resistance processes. in, It is the state transition function. Since the location of the interference source is fixed, only the motion of the aircraft needs to be considered. For radar... Its kinematic model can be described as: The anti-jamming actions of the detector and the power allocation and jamming strategies of the jammer are determined by their respective actions. and The decision is not explicitly specified by the state transition function.
[0033] Game theory payoff: The core objective of the probe is to maximize its combined coverage area. (Time) Cumulative game payoffs within Defined as the comprehensive coverage effectiveness of a radar cluster: in, It is radar At time step The effective coverage area, the specific calculation process of which is given in the radar model later. Within this game theory framework, the objective of the detector is to find an optimal strategy under arbitrary tactics employed by the jammer. To maximize the game's benefits The goal of the jammer is to find an optimal set of strategies regardless of the radar's tactics. To minimize game payoffs ,Right now: Then, a game-theoretic payoff architecture based on a time-varying anisotropic radar detection model.
[0034] In the absence of interference, the radar's maximum detection range is determined by a manually selected false alarm probability. And the farthest effective detection probability To calculate, first, the minimum signal-to-noise ratio, also known as radar sensitivity, is calculated based on the following formula: The effective detection probability at the boundary can be obtained from the radar equations. distance : in This represents the peak power of the radar transmission. and These are the antenna gains for the radar transmitter and receiver, respectively. For the target cross-sectional area, For wavelength, This is the standard noise power of the receiver. Boltzmann's constant, For bandwidth, This refers to the noise temperature of the radar receiver. This refers to system losses.
[0035] Consider that the enemy jammer employs widely used and effective active jamming. During detection, the radar's main lobe remains pointed at the target, while the jammer's main lobe is pointed at the radar, employing a suppressive jamming method. The jammer transmits strong jamming signals to suppress the effective detection signals entering our radar receiver, thereby minimizing the radar's effective detection range and protecting targets within its target area. Given the target point to be detected by the radar... First, the angle between the jammer's main lobe and the radar receiver is calculated using the cosine formula. : Given jammer The interference power applied to it Distance from radar ,radar Received interference power It can be represented as: in, For interference gain, For receiving antenna gain, , This is the antenna directivity coefficient. For radar half-power beamwidth, The overall loss of the interference source, This is the polarization loss of the interference signal.
[0036] In situations with multiple interferences, the radar is at this point. The maximum detection range in the direction can be expressed as: Figure 3 This is a schematic diagram of the detection range under different interference scenarios in the embodiments of the present invention.
[0037] When a radar employs different anti-jamming modes, it can mitigate interference from jammers and increase the radar's detection range. Figure 3As shown. When a radar employs anti-jamming strategies to counter the jamming strategies of a jammer, the effectiveness is often unknown, but it can be estimated through the interaction between the two. For radars with... Radar with anti-jamming strategies and Given the following benefit matrix, consider the jamming machine with various jamming strategies. in, For radar to take the first The first strategy, the jammer adopts the first The benefits gained by the radar under different strategies are discussed. For ease of quantification, different anti-jamming methods are employed by the radar, resulting in different anti-jamming improvement factors. , that is Based on the above formula, the radar's detection range under conditions of multiple jammer interference can be calculated. Finally, the radar... The effective detection range can be expressed as: Step S2: Treat the jamming aircraft cluster as the environment, and model the interaction process between multiple aircraft and the environment as a distributed partially observable Markov decision process, that is, model the multi-agent, multi-system cooperative active game detection problem under jamming confrontation.
[0038] This embodiment treats the jamming aircraft swarm as the environment and models the interaction process between multiple aircraft and the environment as a series of tuples. Defined distributed partially observable Markov decision process. For aircraft clusters, For state space, It is the observation space. For the action space, Let be the state transition probability function. This is a discount factor used to measure the relative importance of immediate and future rewards, with a value between [0,1]. It is the reward function, which directly guides the agent's learning and plays a decisive role in the agent's final policy generation.
[0039] State space: The state space is represented as The status of the aircraft and the state of interference sources in the environment Detailed definitions are provided in step S1.
[0040] Observation space: The observation space of the spacecraft is During electromagnetic spectrum sensing, both the radar and the jammer can obtain information from each other through spectrum analysis. Considering that the frequency of switching countermeasure modes between the radar and the jammer is low, and that the switching process takes a certain amount of time, both sides can obtain the jamming countermeasure mode adopted by the other at the previous decision moment. The following assumptions are made regarding the aircraft in this scenario: (1) The aircraft can communicate with each other, know each other's positions, and maintain the same detection situation map. (2) The aircraft can identify the interference signals it receives and determine the power and type of interference received at the previous moment.
[0041] Therefore, aircraft observation space It consists of two parts: self-information and environmental information. in This refers to the jamming measures and jamming power allocation scheme adopted by the jammer in the previous interaction moment.
[0042] Action space: The action space of the aircraft is ,in Controlled by continuous motion and discrete anti-interference actions It consists of two parts, and detailed definitions can be found in step S1.
[0043] Reward Function: To address the inefficiency and agent behavior stickiness in coupled policy network training, this embodiment proposes a hierarchical incentive mechanism that combines global and individual adaptive rewards. This mechanism utilizes task regions... Discretize into Mutually exclusive grids of varying sizes and maintaining a global access map To quantify the benefits of team collaboration. Individual rewards. It consists of three parts: effective coverage reward, repetitive area adaptive penalty, and anti-interference reward, which together guide the agent's behavior. in, To control the weight of sub-rewards and effectively detect the region It is a collection of the effective detection grids of each aircraft in this operation. Recorded the grid up to the current moment Total number of visits to date It is the base amount for punishment. These are the base values for the penalty exponential growth, parameters that control the severity of the penalty for repeated probes, and determine the relative importance of repeated probes for effective coverage. Additionally, there are rewards for anti-interference strategies. It is a fixed empirical value, obtained by looking up a table based on a pre-defined adversarial benefit matrix. The aircraft swarm obtains its optimal strategy by maximizing the expected reward.
[0044] Step S3: Based on the distributed partially observable Markov decision process, the HyPPO-MACAD multi-agent cooperative anti-interference detection algorithm based on hybrid attention proximal policy optimization is used for training and decision-making to obtain a cooperative policy; the HyPPO-MACAD algorithm introduces a continuous-discrete hybrid branch in the policy network and introduces an attention mechanism in the commentator network.
[0045] This section primarily introduces the Hybrid Attention Multi Agent Proximal Policy Optimization Cooperative anti-jamming detection Algorithm (HyPPO-MACAD). Based on hybrid attention, this improved proximal policy optimization algorithm employs an actor-commentator structure, updating the parameters of the hybrid action space policy network and the hybrid attention value network within a centralized training and distributed execution framework.
[0046] Figure 4 This is a structural diagram of the hybrid action space policy network (left) and the hybrid attention value network (right) in an embodiment of the present invention.
[0047] 1) Hybrid Action Space Policy Network: Aircraft Policy network Parameterization is represented as It consists of two parts: continuous branches and discrete branches, such as Figure 4 As shown, both share two fully connected layers, using tanh as the activation function. Continuous branches consist of a fully connected layer and a Beta distribution layer, with softplus as the activation function in between; discrete branches consist of a fully connected layer and a Category distribution layer, using softmax as the activation function. The input is the observation The outputs are Beta(a,b) and Category(prob), which are obtained by sampling to obtain the pose acceleration and steering angular velocity used for motion control. Discrete anti-jamming actions used for anti-jamming countermeasures : in, It is an abstract feature vector obtained based on observed input. By simultaneously introducing continuous and discrete branches into the network structure, the policy network learns multi-dimensional decisions that combine anti-jamming decision-making and motion control. It can solve the coupled decision-making problem under jamming and adversarial conditions, achieve effective anti-jamming area coverage, and has better adaptability to dynamic adversarial tasks.
[0048] 2) Hybrid Attention Value Network: Under the centralized training and distributed execution framework, all aircraft share a single value network. , parameterized representation as . It consists of two parts: a perceptron composed of three fully connected layers and a hybrid attention module. The hybrid attention module is used to aggregate motion information from other aircraft to improve collaboration efficiency, such as... Figure 4 As shown. The input is the observation vector O = Π of all aircraft. The output is the state score. The hybrid attention module in the observation vector Extract its own motion components and the motion components of other aircraft besides themselves. Based on this, calculate the query matrix (Query) and the key-value pair matching matrix (Key-Value): The attention scores of other aircraft to the current aircraft are calculated using the normalized dot product module: This represents a score indicating the importance of other aircraft relative to the current aircraft. Finally, the observation information... and splicing, inputting into a three-layer perceptron The output yields the state value at the current moment: By introducing a hybrid attention module into the value network, each aircraft obtains the relative importance of other cooperating aircraft to itself, and plans its own motion according to the importance weight. This enables it to guide the local policy learning process based on the global distribution and pose change situation information of the cluster, realize its own optimal path planning and detection strategy, and ultimately reduce detection redundancy and path conflict.
[0049] Figure 5This is a framework diagram of a multi-agent collaborative anti-interference detection algorithm optimized by a hybrid attention proximal strategy in an embodiment of the present invention.
[0050] 3) Training Algorithm Framework: The algorithm uses a centralized training and distributed execution framework, such as... Figure 5 As shown. To ensure that each aircraft can learn global information while simultaneously making motion and anti-interference decisions based on local observation information, the overall process is as described in Algorithm 1. First, set the maximum number of training rounds. Maximum round length Number of aircraft Frequency of policy network parameter updates Initialize the local policy network and global value network And replicate it to the corresponding target network. and Initialize the experience replay pool Then, the environment is initialized, the aircraft cluster interacts with the environment, performs joint actions based on current observations, receives rewards and new observations to obtain experience tuples, which are stored in the experience replay pool. When the experience replay pool reaches a specified size, tuples are sampled from the pool, and the network parameters are updated using the gradient ascent method. Policy Network and evaluation network The corresponding objective function is as follows:
[0051] in, The advantage function, obtained by combining generalized advantage estimation with the hierarchical reward system designed above, guides the training of the agent. Let the objective function of the policy network be... The objective function of the value network, and As a current method for collecting data through real-time interaction between the network and the environment. and The target network completes the parameter updates periodically. 4) Evaluation Metrics: To verify the performance improvement brought by this invention in anti-interference detection tasks, this embodiment provides the area coverage rate within a given time period. and the average number of detections for the detected grid Two metrics are used to quantify the algorithm's performance. The metrics are calculated as follows:
[0052]
[0053] in It is an indicator function that has a value of 1 only when the condition is true, and a value of 0 otherwise.
[0054] Step S4: Based on the trained cooperative strategy, control the aircraft cluster to perform cooperative detection tasks.
[0055] Examples and Performance Verification To verify the effectiveness and superiority of the method proposed in this invention, this section provides a specific software and hardware implementation environment, parameter settings, comparative experiments, and result analysis.
[0056] This embodiment addresses the issues of low radar detection efficiency and easy path entanglement among multiple agents in dynamic and highly contested environments. By combining the aforementioned game-theoretic adversarial model with the hybrid attention proximal strategy optimization algorithm, it provides an implementation process for multi-agent collaborative anti-jamming detection, ensuring efficient target perception and group collaboration in complex time-varying interference scenarios.
[0057] Basic preparations: Hardware environment: The specific implementation was carried out on a computer equipped with an Intel Core i5-1035G1 processor, 8GB of RAM, and a Windows 10 64-bit operating system.
[0058] Software environment: PyCharm Community (Edition 2022.1.1) is used on the computer. The deep network framework is built on PyTorch, and the PyTorch and corresponding CUDA package versions are "torch 1.11.0 +cu113".
[0059] Basic parameters: Task area Defined as Time interval for The longest time a task can be executed for The maximum flight speed of the aircraft for The maximum linear acceleration and the turning angular velocity are respectively and .
[0060] Table 1 shows the resource configuration and initial position distribution information of the aircraft and jamming aircraft.
[0061] Table 1
[0062] Implementation process: Step 1: Construct a specific scenario based on the multi-agent, multi-system collaborative active game detection problem model under interference and adversarial conditions.
[0063] Set the parameters for the game participants: Set the number and initial positions of the probes and the interferers; Establish a time-varying radar detection model: Calculate the detection range based on formulas (9-15) in Example 1, radar parameters, and jammer parameters. The model can be adapted to dynamic jamming environments by real-time data collection and updates from the jammer. Constructing the interference countermeasure benefit matrix: Determining the specific interference countermeasure benefit parameters of formula (16) in Example 1; Determine the game payoff function: With the goal of "maximizing global detection coverage + maximizing interference avoidance success rate", determine the payoff function of the detector in formula (7) of Example 1. With the goal of "interfering with all detectors as much as possible", determine the payoff function of the interfering party in formula (8) of Example 1.
[0064] Step 2: Deploy a multi-agent cooperative anti-interference detection algorithm based on hybrid attention proximal strategy optimization. Design a hybrid action space policy network: Based on formula (22) in Example 1, the observation vector is input into the policy network and encoded into an observation feature vector. Based on the common observation feature vector, formulas (23) and (24) in Example 1 are introduced to process the observation feature vector in parallel, and output continuous actions for controlling the flight trajectory of the aircraft and discrete actions for anti-interference, respectively, to meet the decision variable requirements of the game detection problem.
[0065] Design a hybrid value network: Based on formulas (25-26) in Example 1, and combined with an attention mechanism, the state information of other agents is introduced and fused with the network's own observation information to calculate the attention score between the network and other agents. Through formula (27) in Example 1, information between agents is explicitly fused for calculating state value. The hybrid value network, combined with a reward function, guides the network to obtain lower state values in states where proximity leads to adhesion of the detection area.
[0066] A hierarchical incentive mechanism for global and individual adaptive rewards: The instantaneous reward for each aircraft is calculated using formulas (19-21) in Example 1. For aircraft Its instant rewards Rewards based on the globally effective detection area Individual adaptive repeated detection penalty and anti-interference strategy rewards It consists of three parts. This represents the number of effective grid cells detected by all spacecraft in this operation. The penalty for repeated mesh detection under the action of adaptive coefficients increases exponentially with the number of repeated detections; The improvement gained by adopting anti-jamming strategies for aircraft in response to enemy jamming strategies is a pre-set empirical value.
[0067] Algorithm policy training and update: Input the observation vector into the policy network to generate actions (agent movement acceleration, angular velocity and anti-interference mechanism), the global value network evaluates the value of the generated state of the action, and the policy update amplitude is limited by the clip function to avoid the update amplitude being too large and causing training instability. The specific process can be seen in Algorithm 1. Step 3: Multi-agent cooperative detection execution and evaluation After training, the scenario is first initialized, and then the detection task is executed repeatedly in a loop of "game model perceives environment → policy network generates action → reward mechanism feedback" until the preset detection time is reached to evaluate the algorithm performance. After the task is completed, the area coverage and the number of repeated detections are calculated using formulas (31-32) in Example 1 to evaluate the algorithm performance.
[0068] Comparative verification: To verify the effectiveness and superior performance of the algorithm, the proposed algorithm HyPPO-MACAD is compared with three other benchmark algorithms: MADDPG, MATD3, and MAPPO.
[0069] Table 2 shows the specific hyperparameters of the training algorithm.
[0070] Table 2
[0071] Figure 6 This is a comparison chart of the algorithm training reward curves in an embodiment of the present invention; Training curve comparison: Each algorithm was trained using the same parameters, and the trained model was evaluated every 5000 steps. The resulting reward curves are as follows: Figure 6 The results show that the algorithm in this invention achieves significantly higher rewards than other algorithms, demonstrating the advantages of this algorithmic framework in coupled decision-making problems. Furthermore, the proposed algorithm reduces the adhesion of multi-agent detection regions through hybrid attention and accelerates training based on hierarchical rewards, outperforming MAPPO in both convergence and reward performance.
[0072] Figure 7 Comparison of cooperative detection performance of different algorithms under different agent scales in embodiments of the present invention: Region coverage (left) Average number of detections per covered grid (right) Evaluation Metrics Comparison: Considering different agent sizes, the above-mentioned intelligent algorithms and the bow-shaped search algorithm are used to perform cooperative anti-interference detection tasks within the same time frame. The resulting area coverage and the average number of detections per grid cell in the covered area are as follows: Figure 7As demonstrated, while the bow-shaped search algorithm has a low repetition rate, it suffers from significant missed detection areas under strong interference. The proposed algorithm, based on adaptive coupled decision-making and guided by both a reward layer and a behavior layer, achieves the highest area coverage with minimal detection overlap.
[0073] On the other hand, in order to verify the generalization of the algorithm, we considered task scenarios with different numbers and locations of jammers, and carried out anti-jamming cooperative detection tasks based on the proposed algorithm and other algorithms.
[0074] Table 3 shows the basic configuration information for the scenario.
[0075] Table 3
[0076] The region coverage metrics of the proposed algorithm are shown in Table 4. As can be seen from the table, compared to other advanced algorithms, this algorithm achieves the highest coverage rate in all scenarios, and its number of repeated grid probes is second only to the bow-shaped search algorithm. Although the bow-shaped search algorithm has the lowest number of repeated grid probes, it cannot solve the problem of missed scans caused by interference, thus its region coverage performance is poor. Compared to the bow-shaped search algorithm, the proposed algorithm improves the average coverage rate by 11.46% in three classic scenarios.
[0077] Table 4 shows the results of the generalization experiment comparison.
[0078] Table 4
Claims
1. A multi-agent, multi-system cooperative detection method based on deep reinforcement learning for radar jamming countermeasures, characterized in that, Includes the following steps: Step 1: Construct a dynamic adversarial game model of aircraft clusters and jamming aircraft clusters in the jamming scenario; Step 2: Treat the jamming aircraft cluster as the environment and model the interaction process between the multiple aircraft and the environment as a distributed partially observable Markov decision process. Step 3: Based on the distributed partially observable Markov decision process, the HyPPO-MACAD multi-agent cooperative anti-interference detection algorithm optimized by hybrid attention proximal strategy is used for training and decision-making to obtain the cooperative strategy. The HyPPO-MACAD algorithm introduces a continuous-discrete hybrid branch in the policy network and an attention mechanism in the commentator network. Step 4: Based on the cooperative strategy obtained from training, control the aircraft cluster to perform cooperative detection tasks.
2. The multi-agent, multi-system cooperative detection method based on deep reinforcement learning under radar jamming countermeasures as described in claim 1, characterized in that: in, In step 1, the dynamic adversarial modeling of the aircraft swarm and the jamming aircraft swarm is performed using tuples. The game problem is represented by, where For environmental maps, Gathering for players, and They are state space and action space, respectively. For the policy function, For the environment state transition function, For the sake of game gains.
3. The multi-agent, multi-system cooperative detection method based on deep reinforcement learning under radar jamming countermeasures as described in claim 2, characterized in that: in, The game payoff Based on the calculation of a time-varying anisotropic radar detection model, which is used to calculate the aircraft's position at a point under multiple interference conditions. Detection range in the direction: In the formula, This represents the maximum detection range under interference-free conditions.
4. The multi-agent, multi-system cooperative detection method based on deep reinforcement learning under radar jamming countermeasures as described in claim 1, characterized in that: in, In step 2, the interaction process between multiple aircraft and the environment is modeled as a series of tuples. The defined distributed partially observable Markov decision process, wherein, For aircraft clusters, For state space, It is the observation space. For the action space, Let be the state transition probability function. It is a discount factor. It is a reward function.
5. The multi-agent, multi-system cooperative detection method based on deep reinforcement learning under radar jamming countermeasures as described in claim 4, characterized in that: in, The reward function A hierarchical incentive mechanism combining global and individual adaptive rewards is adopted, with individual rewards... ,in To effectively cover rewards, Adaptive penalty for repetitive regions. This is a reward for preventing interference.
6. The multi-agent, multi-system cooperative detection method based on deep reinforcement learning under radar jamming countermeasures as described in claim 5, characterized in that: in, The formula for calculating the adaptive penalty for the repeated region is: In the formula, For grid Total number of visits.
7. The multi-agent, multi-system cooperative detection method based on deep reinforcement learning for radar jamming countermeasures as described in claim 1. Its features are: in, In step 3, the continuous-discrete hybrid branch of the policy network specifically refers to: the policy network It consists of two parts: a continuous branch and a discrete branch, both of which share two fully connected layers. Continuous branches are constructed using a fully connected layer and a Beta distribution layer; discrete branches are constructed using a fully connected layer and a Category distribution layer.
8. The multi-agent, multi-system cooperative detection method based on deep reinforcement learning under radar jamming countermeasures as described in claim 1, characterized in that: in, In step 3, the attention mechanism in the critic network specifically refers to: the critic network Includes a hybrid attention module for aggregating motion information from other aircraft and calculating the attention scores of other aircraft for the current aircraft. In the formula, , .
9. The multi-agent, multi-system cooperative detection method based on deep reinforcement learning under radar jamming countermeasures as described in claim 1, characterized in that: in, In step 3, the HyPPO-MACAD algorithm employs a centralized training and distributed execution framework, utilizing gradient ascent to optimize the objective function through a proximal strategy. Update policy network parameters: 。 10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 9.