Multi-agent reinforcement learning radar anti-jamming method based on interference capability allocation

By adopting a multi-agent reinforcement learning radar anti-jamming method based on interference capability allocation, the problem of insufficient anti-jamming capability in multi-radar collaborative operation is solved, enabling collaborative detection of radars under conditions of lack of communication, improving the detection success rate, and applicable to radar collaborative anti-jamming scenarios in marine environments.

CN119044899BActive Publication Date: 2025-11-07BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310608519.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-27
Publication Date
2025-11-07
Estimated Expiration
2043-05-27

AI Technical Summary

Technical Problem

In scenarios where multiple radars work together, existing technologies lack sufficient research on anti-jamming capabilities, and continuous communication between radars is difficult to ensure in actual combat, leading to difficulties in collaborative detection and making it hard to guarantee the reliability and accuracy of detection results.

Method used

A multi-agent reinforcement learning radar anti-jamming method based on jamming capability allocation is adopted. Through a centralized training module, jamming identification module, anti-jamming measure selection module, threat assessment module, jamming behavior selection module, and jamming capability allocation module, the radar achieves collaborative anti-jamming among radars, actively changes its working mode to adjust the jamming capability allocation of the jammer, and exposes its weaknesses.

Benefits of technology

It improves the success rate of radar-coordinated anti-jamming detection, is applicable to radar-coordinated anti-jamming scenarios in marine environments, has a wide range of applications, and has good prospects for promotion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119044899B_ABST
    Figure CN119044899B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of radar anti-jamming, and relates to a multi-agent reinforcement learning radar anti-jamming method based on interference capability allocation. The method comprises the following steps: parameter setting; centralized training in cooperation with each radar to learn a suitable strategy; S1 identifies interference behavior, and selects the working mode and anti-jamming measures of each radar according to the identification result and the training strategy; a threat evaluation result is obtained by weighting according to the distance, radar working mode and radar platform speed; after identifying the radar behavior, the behavior of the jammer is selected and the success probability of radar jamming is determined; interference capability is allocated to each radar under the constraint of the radar threat evaluation result and the radar detection success probability; the time consumed in the game process is calculated and updated; when the remaining game time is greater than zero, the game continues, and the process returns to S1; otherwise, the game ends. The method only shares information in training and does not communicate in execution, and is suitable for radar cooperative anti-jamming scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of radar anti-jamming, and relates to a multi-agent reinforcement learning radar anti-jamming method based on interference capability allocation. BACKGROUND

[0002] When a radar detects a non-cooperative target, it may be subjected to various active / passive jamming. For example, in the combat scenario of down-looking detection of sea surface targets by a missile-borne radar, the radar needs to face the complex three-dimensional jamming system on a ship target. The ship platform is large and has strong carrying capacity, and many jamming methods can be released by the ship-borne platform, which has almost no power and size restrictions and has strong jamming capability. The missile-borne radar has limited power and size, and has relatively weak countermeasures. It is very difficult for a single missile-borne radar to complete the detection and anti-jamming task when detecting a ship target, and it is difficult to ensure the reliability and accuracy of the detection results. In order to improve the anti-jamming capability of the missile-borne radar and improve the effect of sea attack, a multi-missile and multi-radar cooperative detection anti-jamming mode can be used to complete the detection task.

[0003] The traditional multi-radar cooperative mode requires continuous communication between radars, which is often difficult to ensure in actual confrontation. The radar has the functions of detection and perception, information processing, and response, and is a typical agent. The detection and anti-jamming capability of the radar can be improved by using the research theory of multi-agent. In the multi-agent cooperative anti-jamming scene, the radar can change the working mode and thus change the threat assessment result of the jammer, mobilize the jammer to allocate the jamming capability, expose the weak points of the jamming capability, and create conditions for the radar to break through the jamming and complete the detection. SUMMARY

[0004] The purpose of the present application is to solve the problems of 1) the current research on multi-radar cooperative work pays insufficient attention to anti-jamming problems, and 2) in actual combat, the continuous communication function between multiple radars is difficult to ensure, so that cooperative detection under the condition of communication loss becomes a major challenge in anti-jamming. The present application provides a multi-agent reinforcement learning radar anti-jamming method based on interference capability allocation.

[0005] In order to achieve the above purpose, the present application adopts the following technical solutions.

[0006] The radar anti-jamming method relies on a radar cooperative anti-jamming system, which includes a centralized training module, a jamming recognition module, an anti-jamming measure selection module, a threat assessment module, a jamming behavior selection module, a jamming capability allocation module, and a win / loss determination module.

[0007] The centralized training module trains the radar cooperation according to prior information before the game process starts. In the confrontation process, the jamming identification module and the anti-jamming measure selection module complete the anti-jamming measures of the radar; the threat evaluation module, the jamming behavior selection module and the jamming capability distribution module complete the jamming measures of the jammer;

[0008] The centralized training module is connected with the jamming identification module, the jamming identification module is connected with the anti-jamming measure selection module, the anti-jamming measure selection module is connected with the threat evaluation module, the threat evaluation module is connected with the jamming behavior selection module, the jamming behavior selection module is connected with the jamming capability distribution module, the jamming capability distribution module is connected with the win-lose determination module, and the win-lose determination module is connected with the jamming identification module and the game result;

[0009] The centralized training module centrally trains each radar, learns the appropriate strategy of changing the working mode and the anti-jamming measure, and each radar can obtain the state information of other radars during the training;

[0010] The jamming identification module allows each radar to identify the jamming behavior and obtain the jamming identification result;

[0011] The anti-jamming measure selection module selects the working mode and the anti-jamming measure of each radar according to the jamming identification result and the training strategy;

[0012] The threat evaluation module collects the radar working mode, distance, speed and other information, obtains the radar threat evaluation result, and provides the jamming capability distribution module of the jammer;

[0013] The jamming behavior selection module identifies the radar behavior, selects the behavior of the jammer according to the softmax, and determines the detection success probability of the radar;

[0014] The jamming capability distribution module distributes the jamming capability of each radar under the constraints of the radar threat evaluation result and the radar detection success probability;

[0015] The win-lose determination module calculates the time required for the behaviors of the jamming identification module, the anti-jamming measure selection module, the threat evaluation module, the jamming behavior selection module and the jamming capability distribution module to take effect, and updates (reduces) the remaining game time with the game process;

[0016] When a round of game proceeds to the win-lose determination module, if the remaining game time is greater than zero, the jamming identification module, the anti-jamming measure selection module, the threat evaluation module, the jamming behavior selection module, the jamming capability distribution module and the win-lose determination module are continued to be executed; if the remaining game time is exactly reduced to zero or smaller, the game is ended, whether the radar completes the anti-jamming is judged, and the game result is output;

[0017] The condition for the radar to complete the anti-jamming is that if all anti-jamming measures of a radar are completed and the radar is in a search mode at the end of the game, while the jamming measures of the jammer are not completed, the radar is considered to complete the anti-jamming task.

[0018] The multi-agent reinforcement learning radar anti-jamming method based on interference capability allocation comprises the following steps:

[0019] S1, parameter setting;

[0020] The parameters in S1 comprise: time required for interference behavior of this round to take effect, time required for the radar to identify interference, time required for the radar to take anti-jamming measures, time required for the jammer to identify radar measures, total interference capability of the jammer, initial distance of the radar, platform speed of the radar, threat weight of the search mode / tracking mode of the radar, interference capability threshold corresponding to the radar being completely jammed, time threshold for the radar being jammed in the tracking mode to switch to the search mode, time threshold for the radar being jammed in the search mode to switch to the tracking mode, and total game time.

[0021] S2, a centralized training module cooperates with each radar to perform centralized training, and learns to obtain suitable strategies for changing the working mode and anti-jamming measures;

[0022] S3, an interference identification module identifies interference behavior and records the identification result;

[0023] S4, an anti-jamming measure selection module selects the working mode and anti-jamming measures of each radar according to the interference identification result and the training strategy;

[0024] S5, a threat evaluation module obtains a threat evaluation result by weighting the distance between the radar and the ship target, the working mode of the radar, and the platform speed of the radar;

[0025] S6, after the interference behavior selection module identifies the radar behavior, the behavior of the jammer is selected and the success probability of the radar being jammed is determined;

[0026] S7, an interference capability allocation module allocates interference capability to each radar under the constraints of the threat evaluation result of the radar and the detection success probability of the radar;

[0027] The interference capability allocated to each radar is achieved by optimizing the target of the jammer, and specifically:

[0028]

[0029]

[0030] wherein, is the optimization target of the jammer, D k is the threat evaluation result of the kth radar, P kThe success probability of the jammer interfering with the kth radar after the jammer allocates interference capability to the kth radar is Pk, J is the total interference capability of the jammer, and θ is the angle between the kth radar and the reference line of the jammer and the radar k The angle between the kth radar and the reference line of the jammer and the radar is θ, A2 is an adjustment coefficient of the peak drop speed, other angles away from the peak are dropped according to a sine function, and only the distribution of the first zero point is considered. The interference capability allocated by the jammer to the kth radar on a single radar is Pk, θ is the angle between the kth radar and the reference line of the jammer and the radar, and A 1k The peak of the interference capability allocated by the jammer to the kth radar is Pk.

[0031] S8, the win or lose determination module calculates the time consumed in the game process and updates the remaining game time.

[0032] The time consumed in the game process includes the time consumed in the identification of the current interference type by the interference identification module, the time consumed in the selection and implementation of the radar measures by the anti-interference measure selection module, the time consumed in the collection and evaluation of information by the threat evaluation module, the time consumed in the selection and implementation of the jammer measures by the interference behavior selection module, and the time consumed in the allocation of the interference capability by the interference capability allocation module.

[0033] The remaining game time is equal to the initial set remaining game time minus the time consumed in the game process.

[0034] S9, when the remaining game time is greater than zero, the game continues, and the process jumps to S3; when the remaining game time is less than or equal to zero, the game ends, and the process jumps to S10.

[0035] S10, the win or lose determination module determines whether the anti-interference measures of a certain radar are all completed and in the search mode, and whether the interference measures of the jammer are not completed, when the game ends.

[0036] The parameters in S1 are parameters set by prior information, and do not change after the interference type and the anti-interference measure type are determined.

[0037] The centralized training in S2 is continuous and dynamic, specifically, each radar is always in the process of switching the working mode in the game, so that the interference capability distribution of the jammer changes continuously, and the multiple radars coordinate the behavior of the jammer, so that the interference capability distribution of the jammer has a weak point, and creates conditions for the radars to complete the detection task.

[0038] The working mode switching specifically includes that the working mode of each radar is actively changed, and the ability of fast switching is retained, the radars change the threat degree to the jammer, so that the jammer passively changes the interference capability distribution.

[0039] S5 said threat assessment results, recorded as D i ; by formula d i is the current missile-borne radar and jammer distance, v i is the radar platform speed, m i is the threat coefficient brought by the radar operating mode.

[0040] S6 said the behavior of selecting the jammer, specifically through the following softmax formula:

[0041]

[0042] Where, P(i) represents the probability of selecting the i-th jamming behavior of the jammer, u i represents the income that the i-th jamming behavior can achieve when facing the current radar behavior.

[0043] S6 said the success probability of the radar being jammed, which is calculated according to the following formula:

[0044]

[0045] Where, P k is the success probability of the jammer jamming the radar, Th1 is the jamming capability threshold corresponding to the radar being completely jammed, A1 is the jamming capability allocated by the jammer on this radar, and A3 is the amplitude adjustment coefficient.

[0046] In the embodiment of the present application, a multi-agent reinforcement learning radar anti-jamming method based on jamming capability allocation is proposed. The radar changes the working mode to change the threat assessment result of the jammer, mobilizes the jamming capability allocation of the jammer to expose the weakness of the jamming capability, and creates conditions for the radar to break through the jamming and complete the detection. The success rate of radar cooperative anti-jamming detection is significantly improved compared with the success rate of each agent optimizing alone. The method can be applied to the scene of radar cooperative anti-jamming in sea environment, has wide application range and good popularization prospect.

[0047] Advantages

[0048] The multi-agent reinforcement learning radar anti-jamming method based on jamming capability allocation proposed in the present application has the following advantages compared with the existing radar cooperative anti-jamming method:

[0049] 1. Each radar only shares information in training, and each radar performs anti-jamming process alone during task execution without communication and cooperative detection behavior;

[0050] 2. The method can be applied to the scene of radar cooperative anti-jamming in sea environment, has wide application range and good popularization prospect. DETAILED DESCRIPTION

[0051] Figure 1 A flowchart of the multi-agent reinforcement learning radar anti-jamming method based on interference capability allocation provided for the present application and embodiments;

[0052] Figure 2 A flowchart of the win-lose determination module timing function of the multi-agent reinforcement learning radar anti-jamming method based on interference capability allocation provided for the present application and embodiments;

[0053] Figure 3 A reward under the scenario of two radar cooperative anti-jamming of the multi-agent reinforcement learning radar anti-jamming method based on interference capability allocation provided for the present application and embodiments;

[0054] Figure 4 A win rate under the scenario of two radar cooperative anti-jamming of the multi-agent reinforcement learning radar anti-jamming method based on interference capability allocation provided for the present application and embodiments;

[0055] Figure 5 A reward under the scenario of four radar cooperative anti-jamming of the multi-agent reinforcement learning radar anti-jamming method based on interference capability allocation provided for the present application and embodiments;

[0056] Figure 6 A win rate under the scenario of four radar cooperative anti-jamming of the multi-agent reinforcement learning radar anti-jamming method based on interference capability allocation provided for the present application and embodiments. DETAILED DESCRIPTION

[0057] Example embodiments will be described more fully hereinafter with reference to the accompanying drawings, in which example embodiments are shown. The example embodiments may, however, be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided as a full and enabling disclosure of the example embodiments, and to fully convey their scope to those skilled in the art. Numbered examples are provided as a full and enabling disclosure of the example embodiments.

[0058] As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.

[0059] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0060] The embodiments can be described with reference to plan views and / or cross-sectional views by virtue of the ideal schematic illustrations of the disclosure. Accordingly, the exemplary illustrations can be modified according to manufacturing techniques and / or tolerances. Therefore, embodiments are not limited to the illustrated embodiments in the drawings, but include modifications based on manufacturing processes. Thus, the zones illustrated in the drawings have schematic properties and the shapes of the zones shown in the drawings illustrate specific shapes of zones of the elements, but are not intended to be limiting.

[0061] To complete the radar cooperative anti-jamming, the embodiment of the disclosure provides a multi-agent reinforcement learning radar cooperative anti-jamming method based on interference capability allocation. The following will be described one by one in combination with the drawings of the embodiments provided by the disclosure.

[0062] Embodiment 1

[0063] As Figure 1 shown, the embodiment of the disclosure provides a multi-agent reinforcement learning radar anti-jamming method based on interference capability allocation, which includes the following steps:

[0064] Step 1, the centralized training module cooperates with each radar to perform centralized training, learns to obtain a suitable strategy for changing the working mode and anti-jamming measures, and each radar can obtain the state information of other radars during training;

[0065] The specific implementation of this step corresponds to S2 in the summary;

[0066] Step 2, the interference identification module identifies the interference behavior currently faced, obtains and records the interference identification result;

[0067] The specific implementation of this step corresponds to S3 in the summary;

[0068] Step 3, the anti-jamming measure selection module selects the working mode and anti-jamming measures of each radar according to the interference identification result obtained in step 2 and in accordance with the training strategy in step 1 for the current interference behavior;

[0069] The specific implementation of this step corresponds to S4 in the summary;

[0070] The training strategy: when a single radar faces the interference measures and the interference capability allocation of the jammer, it judges according to the interference identification result and the detection success probability, and the state such as the distance from the target, the radar platform speed, etc., selects the appropriate anti-jamming measures and working mode according to the multi-radar cooperative experience obtained by training, mobilizes the interference capability allocation of the jammer, makes the interference capability allocation of the jammer repeatedly change among the radars, so as to appear an interference capability gap, and increase the radar detection anti-jamming probability.

[0071] Step 4, the threat assessment module collects radar working mode, distance, speed and other information to provide basis for subsequent decision-making. And according to the collected data information, the threat degree of each radar to the ship target is evaluated, and the radar threat assessment result is obtained and provided to the jammer;

[0072] The specific implementation of this step corresponds to S5 in the summary of the invention;

[0073] Step 5, after identifying the radar behavior according to the data information in step 4, the jammer behavior is selected according to the softmax, and the radar detection success probability is determined;

[0074] The specific implementation of this step corresponds to S6 in the summary of the invention;

[0075] Step 6, the interference capability distribution module makes a decision, under the constraints of the radar threat assessment result and the radar detection success probability, the interference capability is distributed to each radar, that is, the interference capability is distributed to the direction where each radar is located, and the value of the interference capability of the ship jammer in the direction corresponding to the radar is determined;

[0076] The specific implementation of this step corresponds to S7 in the summary of the invention;

[0077] Step 7, the win-lose determination module has a timing function and a win-lose determination function. The specific timing function process is as shown in Figure 2 When the remaining game time is greater than zero, the game continues, and jumps to step 2; when the remaining game time is less than or equal to zero, the game ends, and if there is a radar whose anti-jamming measures are all completed and is in the search mode, and the jammer's interference measures are not completed, it is judged that the radar completes the anti-jamming task;

[0078] The specific implementation of this step corresponds to S8-S10 in the summary of the invention;

[0079] It should be noted that the specific timing function process is as shown in Figure 2 According to the game time, after each round of game ends, the radar identification interference time, the radar anti-jamming measure time, the jammer radar identification time and the jammer interference measure time are subtracted, and the remaining game time is updated, including the following steps:

[0080] Step 1, set the game time;

[0081] Step 2, time the game process to obtain the game process time, specifically: in each round of game, the sum of the radar identification interference time, the radar anti-jamming measure time, the jammer radar identification time and the jammer interference measure time;

[0082] Step 3, judge whether the remaining game time is positive, if yes, update the game time and jump to step 2, continue to time the game process; otherwise, the game ends, and the timing stops.

[0083] The game time refers to the total game time length set artificially in simulation, and refers to the total game time length calculated through a series of parameters such as the distance between the bullet and the target and the relative speed between the bullet and the target in actual combat.

[0084] The remaining game time refers to the game time remaining after each round of game ends, which is equal to the game time minus the game process time.

[0085] The time required for the radar to identify the interference refers to the time required for the radar to receive the interference signal and judge the interference type;

[0086] The time required for the radar to take anti-interference measures refers to the time required for the radar to select a targeted anti-interference method according to the current interference type and take effect;

[0087] The time required for the jammer to identify the radar refers to the time required for the jammer to collect radar information and calculate the radar threat assessment;

[0088] The time required for the jammer to take interference measures refers to the time required for the jammer to select interference behavior and allocate interference capability to each radar and take effect;

[0089] In the radar cooperative anti-interference method, the clutter, noise and other factors are not considered, only the behaviors that can be controlled by the radar and the jammer are considered.

[0090] The multi-radar cooperative detection anti-interference method constructs a scene of multi-radar and jammer confrontation, under the condition of lacking communication between radars, uses multi-agent reinforcement learning to optimize the behavior measures of each radar, improves the detection anti-interference ability, and improves the success probability of the radar completing the detection task. In specific implementation, the proposed algorithm is compared with other three algorithms, including DDPG algorithm (each agent uses DDPG algorithm for optimization in training and execution), conventional strategy (each agent uses the strategy with the best anti-interference effect for the current interference) and random strategy (each agent randomly selects its strategy). In order to show the advantages of the proposed method in selecting multi-agent cooperative strategy, two radars and four radars are used for simulation experiments.

[0091] Two-radar simulation experiment: in the experiment, it is assumed that the radar has 5 strategies to choose from, and the jammer has 4 strategies to choose from. The initial distance of the missile-borne radar is 80 kilometers, and the radar platform speed is 1 kilometer per second. The game ends when the radar platform hits the ship target, the radar falls into the sea, the radar is not enough to complete the maneuver, etc. The anti-interference success probability of the radar in simulation for different interference styles is set as follows:

[0092]

[0093] From the foregoing analysis, since the identification and action of the radar and the jammer need to consume time, the strategy of always choosing the highest benefit action for the current jammer is not necessarily the optimal strategy overall. For example, the radar takes the first anti-jamming action against the first jammer, which can achieve good results at present. After the jammer changes the action in the subsequent game process, the anti-jamming benefit will decrease when the radar re-performs the jammer identification period. In this case, taking other actions can bring greater benefits, which requires the radar to optimize the cooperative strategy.

[0094] The cooperation between the two radars mainly reflects the cooperative change of the working mode, and the difference in distance and speed is not much. The threat difference of the ship jammer is mainly in the working mode. In the multi-agent cooperation, the missile-borne radar can break through the jamming and detect the target when other radars are in the tracking state and can withstand more jamming ability of the jammer.

[0095] In the jamming game confrontation model of multiple radars, the overall game winning rate is taken as an index to evaluate the performance of the MADDPG (Multi-Agent Deep Deterministic Policy Gradient) algorithm based on jamming ability allocation, and compared with the single-agent DDPG (Deep Deterministic Policy Gradient) algorithm, the regular strategy and the random strategy. The parameters involved in the simulation are as follows in Table 1:

[0096] Table 1 Simulation parameter table

[0097]

[0098]

[0099] The rewards and winning rates in the two-radar cooperative anti-jamming scenario are shown in Figure 3 and Figure 4 The experimental results show that the two missile-borne radars can actively control the working mode and anti-jamming action to create a weak point of the jammer and break through the jamming defense to complete the cooperative detection task.

[0100] Four-radar experiment: In this simulation experiment, the jamming ability threshold (Th1) at which the radar is completely jammed is set to 0.4, and the other parameters are consistent with those in the two-radar experiment. The jamming ability threshold (Th1) at which the radar is completely jammed should be a fixed value in practice and does not change with the number of cooperative radars. This experiment is only to verify the effectiveness of the MADDPG algorithm based on jamming ability allocation for different numbers of radars, and does not have a comparative effect with the two-radar experiment.

[0101] The rewards and winning rates in the four-radar cooperative anti-jamming scenario are shown in Figure 5and Figure 6 As shown. By Figure 5 It is evident that in a four-agent scenario, the benefits of multi-agent collaboration are significant. Due to environmental changes caused by the actions of other agents, the DDPG algorithm still fails to learn a stable policy, resulting in a lower reward value compared to the policy learned by the MADDPG algorithm. The MADDPG algorithm, based on interference capability allocation, learns stable policies through concentrated training on the states and policies of other agents, achieving faster convergence and higher rewards. Other concentrated policies yield lower rewards, leaving the radar consistently at a disadvantage in adversarial situations. Figure 6 It can be seen that in a scenario where four agents cooperate, the strategy using the MADDPG algorithm has a significantly higher win rate than the DDPG algorithm optimized by each agent individually, the conventional strategy, and the random strategy. Although the win rate fluctuates due to the influence of randomness, the MADDPG algorithm based on the allocation of interference capabilities can bring a higher average win rate, reaching 70% at the end of the game, which is higher than the results of individual optimization by each agent, as well as the results of the random strategy and the conventional strategy.

[0102] The reward stability of the conventional and stochastic policies indicates that the policies are stable because they do not take any optimization measures. The results of the two experiments show that the proposed method can learn a stable policy using the MADDPG algorithm based on perturbation capability allocation. During training, an increase in reward value can be observed at the beginning of training, followed by stabilization after learning the optimal policy.

[0103] Training with the MADDPG algorithm based on jamming capability allocation yields the optimal multi-agent cooperative strategy, maximizing the overall radar reward in the MDP. Experiments show that the optimal strategy trained using the MADDPG algorithm enables radar cooperation and generates a detection synergy effect, thus better completing the detection task. The jamming capability allocation model in the model indicates that tracking mode radars attract more jamming capabilities. Experimental results show that the strategy obtained using the MADDPG algorithm based on jamming capability allocation can effectively weaken radar jamming in a specific direction, enabling other radars to complete the detection task more effectively. Multiple radars actively change their operating modes, altering their threat priority to jammers and adjusting the jammer's jamming capability allocation, passively exposing its jamming capability gaps, thereby ensuring that at least one radar in the network can complete the detection task.

[0104] The multi-radar cooperative detection anti-jamming based on multi-agent reinforcement learning has great similarity with multi-missile cooperative attack of ship targets. The multi-missile cooperative attack generally launches from different locations, flies according to the planned path, and achieves the effect of attacking the target from different angles at the same time. The whole cooperative attack process defaults to hit the target as long as the preset scene of attacking the target from different angles at the same time is realized. During the execution of the task, each missile-borne radar independently performs detection anti-jamming without communication and cooperative detection behavior, which has great difficulty. The radar anti-jamming strategy optimization method based on multi-agent reinforcement learning is to train the radar behavior setting in the flight process in advance. These behaviors can realize cooperation in the flight process by changing the working mode, but the essence is still a kind of program pre-control behavior. Unlike the general program pre-control behavior, the general program behavior setting is relatively simple, without considering reasonable jammer modeling and complex interaction between radar and jammer. In many confrontation models, the decision of jamming is simply set as randomly extracting a suitable jamming behavior at a certain missile-target distance, which is not high in intelligence, so it is difficult to realize cooperation. In the radar anti-jamming strategy optimization method based on multi-agent reinforcement learning, the decision process of the jammer is more reasonably set, the radar decision behavior is better trained, the detection effect is higher, and the multi-radar occupies an advantage in the confrontation with the jammer.

[0105] In the embodiment of the application, each radar acquires cooperative tacit understanding in centralized training, intelligently adjusts the working mode in distributed execution, actively changes the threat priority of the radar, mobilizes the allocation of jamming capability of the jammer, makes the jamming capability of the jammer passively produce a concave point, so as to ensure that at least one radar in the networked radar can complete the detection task. The success rate of radar cooperative anti-jamming detection is significantly improved compared with the success rate of each agent optimizing alone. The method can be applied to the radar cooperative anti-jamming scene in the sea environment, has a wide application range and good popularization prospect.

[0106] The above is only an optional implementation mode of part of the implementation scenarios of the application, and it should be pointed out that, for ordinary technical personnel in the technical field, other similar implementation means based on the technical idea of the application without departing from the technical concept of the application also belong to the protection scope of the embodiments of the application.

Claims

1. A method for radar anti-jamming based on multi-agent reinforcement learning with interference capability allocation, characterized in that, The method comprises the following steps: S1, parameter setting; The parameters in S1 comprise: time required for the current interference behavior to take effect, time required for the radar to identify the interference, time required for the radar to take anti-interference measures, time required for the jammer to identify the radar measures, total interference capability of the jammer, initial distance of the radar, platform speed of the radar, threat weight of the radar in search mode / tracking mode, interference capability threshold corresponding to the radar being completely interfered, time threshold for the radar in tracking mode to be interfered to switch to search mode, time threshold for the radar in search mode to switch to tracking mode, and total game time; S2, a centralized training module cooperates with each radar to perform centralized training, and learns to obtain a suitable strategy for changing the working mode and the anti-interference measures; S3, an interference identification module identifies the interference behavior and records the identification result; S4, an anti-interference measure selection module selects the working mode and the anti-interference measures of each radar according to the interference identification result and the training strategy; S5, a threat evaluation module obtains a threat evaluation result according to the distance between the radar and the ship target, the working mode of the radar, and the platform speed of the radar; S6, after the radar behavior is identified by the interference behavior selection module, the behavior of the jammer is selected, and the success probability of the radar being interfered is determined; S7, an interference capability distribution module distributes the interference capability to each radar under the constraint of the threat evaluation result of the radar and the radar detection success probability; The interference capability distribution to each radar is achieved by optimizing the jammer target, and specifically: wherein, is the optimization target of the jammer, is the threat assessment result of the first radar, is the success probability of the jammer jamming the first radar after the jammer allocates jamming capability to the radar, is the total jamming capability of the jammer, is the angle of the first radar deviating from the reference line between the jammer and the radar, the reference line between the jammer and the radar being determined by the jammer selecting a certain radar and connecting with the radar; is the adjustment coefficient of the peak drop speed, other angles away from the peak drop according to the sinc function and only considering the distribution of the first zero point; represents the jamming capability allocated by the jammer of the first radar to the single radar, is the angle deviating from the reference line between the radar platform and the jammer, is the peak of the jamming capability allocated by the jammer to the first radar. S8, a win-loss determination module calculates the consumed time of the game process and updates the remaining game time; The consumed time of the game process comprises: the identification time of the current interference type by the interference identification module, the time for the anti-interference measure selection module to select the radar measures and take effect, the time for the threat evaluation module to collect information and evaluate the radar threat, the time for the interference behavior selection module to select the jammer measures and take effect, and the time required for the interference capability distribution module to distribute the interference capability; The remaining game time is equal to the initial set remaining game time minus the consumed time of the game process; S9, when the remaining game time is greater than zero, the game continues, and jumps to S3; when the remaining game time is less than or equal to zero, the game ends, and jumps to S10; S10, the win-loss determination module determines whether the radar completes the anti-interference task when the game ends, that is, whether there is a radar whose anti-interference measures are all completed and is in search mode, and whether the jammer measures are not completed.

2. The multi-agent reinforcement learning radar anti-jamming method according to claim 1, characterized in that, The parameters in S1 are parameters set by prior information, which do not change after the type of interference and the type of anti-interference measures are determined.

3. The multi-agent reinforcement learning radar anti-jamming method according to claim 1, characterized in that, The centralized training in S2 is continuous and dynamic, specifically: each radar is always in the working mode switching process in the game, so that the interference capability distribution of the jammer changes continuously; multiple radars cooperate to mobilize the behavior of the jammer, so that the interference capability distribution of the jammer has a weak point, which creates conditions for the radar to complete the detection task.

4. The multi-agent reinforcement learning radar anti-jamming method according to claim 3, characterized in that, The working mode switching is specifically: by actively changing the working mode of each radar and retaining the ability to quickly switch, the radar changes the threat degree to the jammer, so that the jammer passively changes the interference capability distribution.

5. The multi-agent reinforcement learning radar anti-jamming method according to claim 1, characterized in that, S5 the threat assessment result, recorded as ; calculated by formula , is the current missile-borne radar and jammer distance, is the radar platform speed, is the threat coefficient brought by the radar operating mode.

6. The multi-agent reinforcement learning radar anti-jamming method according to claim 1, characterized in that, The behavior of the jammer is selected according to the following softmax formula: wherein, represents the probability that the jammer selects the jth jamming behavior, represents the payoff that the jth jamming behavior can achieve when facing the current radar behavior.

7. The multi-agent reinforcement learning radar anti-jamming method according to claim 1, characterized in that, The success probability of the radar being jammed is calculated according to the following formula: wherein, P is the success probability of the jammer to jam the radar, P is the success probability of the jammer to jam the radar, P is the success probability of the jammer to jam the radar, is the amplitude adjustment coefficient.

8. A multi-agent reinforcement learning radar countermeasure system based on interference capability allocation, characterized in that, The system comprises a centralized training module, a jamming identification module, an anti-jamming measure selection module, a threat evaluation module, a jamming behavior selection module, a jamming capability distribution module and a win-lose judgment module. The centralized training module trains the radar cooperation according to prior information before the game process starts; during the confrontation process, the jamming identification module and the anti-jamming measure selection module complete the anti-jamming measures of the radar; the threat evaluation module, the jamming behavior selection module and the jamming capability distribution module complete the jamming measures of the jammer. The centralized training module is connected with the jamming identification module, the jamming identification module is connected with the anti-jamming measure selection module, the anti-jamming measure selection module is connected with the threat evaluation module, the threat evaluation module is connected with the jamming behavior selection module, the jamming behavior selection module is connected with the jamming capability distribution module, the jamming capability distribution module is connected with the win-lose judgment module, and the win-lose judgment module is connected with the jamming identification module and the game result. The centralized training module centrally trains each radar, learns the appropriate strategy of changing the working mode and the anti-jamming measure, and the state information of other radars can be obtained by each radar during the training; The jamming identification module allows each radar to identify the jamming behavior and obtain the jamming identification result; The anti-jamming measure selection module selects the working mode and the anti-jamming measure of each radar according to the jamming identification result and the training strategy; The threat evaluation module collects the radar working mode, distance, speed and other information, obtains the radar threat evaluation result, and provides it to the jamming capability distribution module of the jammer; The jamming behavior selection module identifies the radar behavior, selects the behavior of the jammer according to the softmax, and judges the detection success probability of the radar; The jamming capability distribution module distributes the jamming capability of each radar under the constraints of the radar threat evaluation result and the radar detection success probability; The win-lose judgment module calculates the time required for the behaviors of the jamming identification module, the anti-jamming measure selection module, the threat evaluation module, the jamming behavior selection module and the jamming capability distribution module to take effect, and updates the remaining game time with the game process; When a round of game is carried out to the win-lose judgment module, if the remaining game time is greater than zero, the jamming identification module, the anti-jamming measure selection module, the threat evaluation module, the jamming behavior selection module, the jamming capability distribution module and the win-lose judgment module are continued to be executed; if the remaining game time is exactly reduced to zero or smaller, the game is ended, and whether the radar completes the anti-jamming is judged and the game result is outputted; The condition for the radar to complete the anti-jamming is that if the anti-jamming measures of a radar are all completed and the radar is in the search mode at the end of the game, and the jamming measures of the jammer are not completed, it is considered that the radar completes the anti-jamming task.

Citation Information

Patent Citations

  • Radar countermeasure strategy modeling and simulation method based on zero-sum game

    CN112651181A

  • Radar jammer game strategy acquisition method based on deep Q network

    CN115113146A