A method for agile frequency multi-radar cooperative jamming resistance based on reinforcement learning

By modeling the radar cooperative anti-jamming process as a generalized Markov decision process and utilizing the parallel multi-agent Q-learning algorithm, the multi-radar system collaboratively selects carrier frequency strategies, solving the problem of insufficient anti-frequency sweeping interference capability of a single radar and achieving higher signal-to-noise ratio and frequency coverage.

CN116125397BActive Publication Date: 2026-02-03UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310066921.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-13
Publication Date
2026-02-03
Estimated Expiration
2043-01-13

AI Technical Summary

Technical Problem

When faced with strong interference, existing frequency-agile radars cannot effectively resist frequency-sweeping interference using random frequency hopping or intelligent frequency modulation methods of a single radar, and there is a lack of research on multi-radar cooperative anti-interference based on reinforcement learning.

Method used

Each radar is treated as an agent, and a generalized Markov decision process is established. The parallel multi-agent Q-learning algorithm (PMAQL) is used to solve the radar carrier frequency selection strategy. Through multi-radar cooperative anti-jamming, interference frequency bands are avoided.

Benefits of technology

It effectively suppresses frequency sweep interference, improves the signal-to-noise ratio (SINR) of the radar, reduces the interference level of individual radars, expands the frequency coverage range, and enhances anti-jamming performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116125397B_ABST
    Figure CN116125397B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on reinforcement learning's agile frequency multi-radar cooperative jamming method, first according to phased array radar signal processing flow, target echo model and interference signal model are established, each radar is regarded as an agent, and radar cooperative jamming process is modeled as generalized Markov decision process, state value function is obtained, the problem is solved using the proposed parallel multi-agent Q learning algorithm, finally radar carrier frequency selection strategy can be obtained.The method of the application can select the appropriate frequency band of radar according to the current frequency band of interference, so as to avoid the possible frequency band of interference at next time, by modeling the multi-radar cooperative jamming process as a generalized Markov decision process, the radar carrier frequency is regarded as an action, the interference carrier frequency is regarded as a state, the SINR is used as a reward function, the radars cooperate with each other, reduce the interference degree of a single radar, and effectively suppress the frequency sweeping interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of radar anti-jamming technology, specifically relating to a frequency-agile multi-radar cooperative anti-jamming method based on reinforcement learning. Background Technology

[0002] Frequency-agile radar, as a typical electronic countermeasures device, has been widely used in various electronic warfare scenarios. It achieves anti-jamming by rapidly changing the signal carrier frequency within, between, or between pulse groups, and features long detection range, high angle measurement accuracy, and strong anti-jamming capability. Frequency sweeping jamming, as a commonly used electronic countermeasures technique, reduces radar resolution and detection performance by dynamically scanning radar frequencies. With the enhancement of jamming capabilities and the development of cognitive jamming, the random or pseudo-random frequency hopping strategies of frequency-agile radar can no longer achieve effective anti-jamming performance. Therefore, researching a new and effective anti-jamming technology is an urgent need for future electronic warfare applications.

[0003] To adapt to complex and ever-changing adversarial environments, radar systems need to possess autonomous learning capabilities. Therefore, intelligent anti-jamming technology based on reinforcement learning has been proposed, whereby the radar learns jamming strategies through interaction with the external environment. The radar anti-jamming process is modeled as a Markov Decision Process (MDP), and the optimal strategy is solved using a reinforcement learning algorithm. The literature "Dong Shuxian, Wu Yaojun, Fang Wen, Quan Yinghui. Frequency-agile radar combined with fuzzy C-means anti-intermittent sampling jamming. Journal of Radar, 2022, 11(02): 289-300." proposes an anti-ISRJ method based on frequency-agile radar combined with fuzzy C-means (FCM). A radar transmit waveform with inter-pulse frequency random agility is designed, which can effectively counter ISRJ interference. The paper "L. Kang, J. Bo, L. Hongwei and L. Siyuan. Reinforcement Learning based Anti-jamming Frequency Hopping Strategies Design for Cognitive Radar. 2018 IEEE International Conference on Signal Processing, Communications and Computing (ICSPCC), Qingdao, China, 2018, pp. 1-5" proposes a single-radar frequency hopping strategy design method based on reinforcement learning to overcome the limitations of using random frequency hopping to combat unknown interference. The radar learns the jammer's strategy through interaction with the environment and takes optimal action to obtain higher rewards. One method uses random frequency hopping to combat intermittent sampling interference with a single-station radar under known interference information, while the other focuses on reinforcement learning-based anti-sweeping interference research for a single-station radar. However, when the interference capability is strong, a single radar cannot achieve good anti-jamming capability using random frequency hopping or intelligent frequency modulation. Therefore, multi-radar cooperative anti-jamming technology can be adopted. Currently, there is limited research on multi-radar cooperative anti-sweeping interference based on reinforcement learning; therefore, further research is necessary. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention proposes a reinforcement learning-based method for cooperative anti-jamming using multiple frequency-agile radars.

[0005] The technical solution adopted in this invention is: a frequency-agile multi-radar cooperative anti-jamming method based on reinforcement learning, the specific steps of which are as follows:

[0006] Step S1: Establish the target echo model and interference signal model according to the phased array radar signal processing flow;

[0007] Step S2: Treat each radar as an intelligent agent, model the radar cooperative anti-jamming process as a generalized Markov decision process, and obtain the state value function.

[0008] Step S3: Solve the generalized Markov decision problem in step S2 using the parallel multi-agent Q-learning algorithm to obtain the radar carrier frequency selection strategy.

[0009] Furthermore, step S1 is specifically as follows:

[0010] The radar system is set to exist. There are 10 radar nodes, and the distance between any two radars is negligible compared to the distance between the radar and the target.

[0011] The transmitted signal from each radar node undergoes carrier frequency agility between pulses, and the number of pulses in a pulse train is... . Indicates radar node The set of carrier frequencies that can be selected for the transmitted pulse.

[0012] in, The first in Each element is represented as , This represents a fixed frequency step value, set relative to the radar bandwidth. Same size This indicates the number of selectable carrier frequencies for a radar pulse; and it is set that any two selectable carrier frequencies do not overlap, and the radar node... Transmitted pulse train signal Represented as:

[0013] (1)

[0014] in, Indicates time, Indicates the pulse repetition period. Indicates radar node The The carrier frequency of each pulse This represents a linear frequency modulated signal with a unit amplitude.

[0015] Radar Node The transmitted signal illuminates the target, generating a target echo, which then triggers the radar node. In the The echo signal after being affected by interference and noise at each pulse point Represented as:

[0016] (2)

[0017] in, Indicates time delay. Indicates radar node With the goal in the Distance per pulse, Indicates the speed of propagation of electromagnetic waves. Indicates Doppler frequency shift, Let the power be zero mean. Gaussian white noise, Indicates unit interference signal, This indicates the change in amplitude caused by the propagation of the interference signal in space; Let represent a parameter that incorporates propagation effects and target scattering; its amplitude is expressed as:

[0018] (3)

[0019] in, Indicates radar node The The received power of each pulse at the radar receiver Indicates the transmission power. These represent the transmit antenna gain and the receive antenna gain, respectively. Indicates the signal wavelength. This represents the target's cross section (RCS).

[0020] The jammer is configured to use a frequency sweeping strategy to interfere with radar signals. This indicates that the interference signal can be selected from a set of carrier frequencies.

[0021] in, This indicates the number of selectable carrier frequencies for the interference signal. The first in Each element is represented as , This represents a fixed interference frequency step value, where the interference carrier frequency is randomly or according to a certain strategy from... Select the appropriate option to set the frequency hopping range of the interference to cover all possible frequency bands of the radar system; set the interference pulse and radar pulse to be synchronized in time, and the amplitude of the interference signal received by the radar to be:

[0022] (4)

[0023] in, This indicates the interference power at the radar receiver. Indicates the jamming transmission power. Indicates the interfering carrier frequency. This indicates the gain of the jamming transmitting antenna, and the jammer's effect on the radar node. Interference probability for:

[0024] (5)

[0025] in, and They represent the first time. The interference airborne frequency is [frequency value] per pulse. Intermediate frequency bandwidth and radar nodes carrier frequency is The intermediate frequency bandwidth.

[0026] radar node In the The signal-to-interference-plus-noise ratio (SINR) of a pulse is expressed as:

[0027] (6)

[0028] Furthermore, step S2 is specifically as follows:

[0029] Step S21: Establish a generalized Markov decision process;

[0030] Each radar node in the radar system is considered an intelligent agent, and the radar cooperative anti-jamming process is modeled as a generalized Markov decision process, consisting of a quintuple. The specific definition is as follows:

[0031] (1) Intelligent agent set All intelligent frequency-agile radars constitute this intelligent entity set. ;

[0032] (2) Action set The selectable carrier frequencies of all radars constitute the action set. Any two radar nodes and of and They do not intersect; radar nodes In the Time step (i.e., the first) (pulse) action The launch of the The carrier frequency representation of each pulse, i.e. ;

[0033] (3) State set The state set consists of all selectable carrier frequencies of the jammer. Radar node In the state of time step Represented by the carrier frequency of the interference, i.e. ;

[0034] (4) State transition probability : Indicates radar node From state Execute action at time The state transitions to The transition probability is expressed as:

[0035] (7)

[0036] in, It is considered unknown.

[0037] (5) Reward Set The reward set consists of the SINR of each pulse from all radars. Radar node In the The reward at each time step is represented as , It is obtained from equation (6).

[0038] Step S22: Obtain the state-action value function;

[0039] Set the optimal policy for each agent for:

[0040] (8)

[0041] in, The state-action value function, or Q-function, is defined as follows:

[0042] (9)

[0043] in, This indicates the calculation of mathematical expectation. Indicates radar node In the Rewards at each time step Indicates the discount rate. Indicates the first After the first time step Each time step This represents the weighted value obtained from the discount rate. Indicates the first The discounted return at the time step is determined by future returns. We obtain the weighted sum.

[0044] Furthermore, in step S3, the PMAQL algorithm solution process is as follows:

[0045] Step S31, let Initialize action set State set Discount rate Learning rate ;

[0046] Among them, learning rate Its value decreases as the iteration proceeds.

[0047] Step S32, Hypothesis: Co-training There are n coherent processing intervals CPI, each CPI is considered as one iteration, and in the nth iteration... At the start of the next iteration, the initial actions are initialized randomly. Interference carrier frequency Obtain the initial state ;

[0048] Step S33: For each pulse in each CPI ,according to Greedy strategy selects action , ;

[0049] Step S34: Sensing the interfering carrier frequency from the environment. ;

[0050] Step S35, Integration and The reward was calculated. , get the state ;

[0051] Step S36: Update the intelligent agent of function ;

[0052] (10)

[0053] in, They represent .

[0054] Step S37: Update Status ,when If the condition is met, return to step 33; otherwise, proceed to step S38.

[0055] Step S38, let ,like Otherwise, proceed to step S39;

[0056] Step S39: Output the final parallel output. Function, denoted as ;

[0057] (11)

[0058] in, This indicates the action selection of multiple agents. Each time the radar pulse selects a carrier frequency, it searches for the detected interference carrier frequency. This allows us to obtain the optimal carrier frequency and achieve anti-interference.

[0059] The beneficial effects of this invention are as follows: First, based on the phased array radar signal processing flow, a target echo model and an interference signal model are established. Each radar is considered an intelligent agent, and the radar cooperative anti-jamming process is modeled as a generalized Markov decision process to obtain the state value function. The proposed parallel multi-agent Q-learning algorithm is then used to solve the problem, ultimately yielding a radar carrier frequency selection strategy. This method can select a suitable radar frequency band based on the current frequency band of the interference, thereby avoiding the possible frequency band of the interference at the next moment. By modeling the multi-radar cooperative anti-jamming process as a generalized Markov decision process, with the radar carrier frequency considered as an action, the interference carrier frequency as a state, and SINR as the reward function, the radars cooperate with each other, reducing the interference level on individual radars and effectively suppressing frequency sweeping interference. Attached Figure Description

[0060] Figure 1 This is a flowchart of a frequency-agile multi-radar cooperative anti-jamming method based on reinforcement learning according to the present invention.

[0061] Figure 2 This is a schematic diagram of multi-radar cooperative anti-jamming in an embodiment of the present invention.

[0062] Figure 3 This is a flowchart of the PMAQL algorithm solution in an embodiment of the present invention.

[0063] Figure 4 The figure shows the results of simulating the number of interference pulses under various methods in the embodiments of the present invention.

[0064] Figure 5 The figures show the SINR results under various simulation methods in the embodiments of the present invention.

[0065] Figure 6 The image shows the interference and radar carrier frequency results of the final iteration of the PMAQL algorithm in this embodiment of the invention. Detailed Implementation

[0066] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0067] like Figure 1 The flowchart of a frequency-agile multi-radar cooperative anti-jamming method based on reinforcement learning according to the present invention is shown below. The specific steps are as follows:

[0068] Step S1: Establish the target echo model and interference signal model according to the phased array radar signal processing flow;

[0069] Step S2: Treat each radar as an intelligent agent, model the radar cooperative anti-jamming process as a generalized Markov decision process, and obtain the state value function.

[0070] Step S3: Solve the generalized Markov decision problem in step S2 using the Parallel Multi-Agent Q-Learning (PMAQL) algorithm to obtain the radar carrier frequency selection strategy.

[0071] In this embodiment, step S1 is specifically as follows:

[0072] The radar system is set to exist. There are 10 radar nodes, and the distance between any two radars is negligible compared to the distance between the radar and the target.

[0073] The transmitted signal from each radar node undergoes carrier frequency agility between pulses, and the number of pulses in a pulse train is... . Indicates radar node The set of carrier frequencies that can be selected for the transmitted pulse.

[0074] in, The first in Each element is represented as , This represents a fixed frequency step value, set relative to the radar bandwidth. Same size This indicates the number of selectable carrier frequencies for the radar pulse; and it is set that any two selectable carrier frequencies do not overlap, as shown in the scene diagram. Figure 2 As shown, radar node Transmitted pulse train signal Represented as:

[0075] (12)

[0076] in, Indicates time, This indicates the pulse repetition time (PRT). Indicates radar node The The carrier frequency of each pulse A linear frequency modulated signal with a unit amplitude is represented as:

[0077] (13)

[0078] in, Indicates the pulse width. Indicates the frequency modulation slope. Represents a matrix function.

[0079] Radar Node The transmitted signal illuminates the target, generating a target echo, which then triggers the radar node. In the The echo signal after being affected by interference and noise at each pulse point Represented as:

[0080] (14)

[0081] in, Indicates time delay. Indicates radar node With the goal in the Distance per pulse, Indicates the speed of propagation of electromagnetic waves. Indicates Doppler frequency shift, Let the power be zero mean. Gaussian white noise, Indicates unit interference signal, This indicates the change in amplitude caused by the propagation of the interference signal in space; Let represent a parameter that incorporates propagation effects and target scattering; its amplitude is expressed as:

[0082] (15)

[0083] in, Indicates radar node The The received power of each pulse at the radar receiver Indicates the transmission power. These represent the transmit antenna gain and the receive antenna gain, respectively. Indicates the signal wavelength. This represents the target scattering cross section (RCS).

[0084] The jammer is configured to use a frequency sweeping strategy to interfere with radar signals. This indicates that the interference signal can be selected from a set of carrier frequencies.

[0085] in, This indicates the number of selectable carrier frequencies for the interference signal. The first in Each element is represented as , This represents a fixed interference frequency step value, where the interference carrier frequency is randomly or according to a certain strategy from... Select from the options to set the frequency hopping range of the interference to cover all possible frequency bands of the radar system; set the interference to... The probability is step frequency interference, with If the frequency point is randomly selected with a probability, then the transition probability matrix of the interference is... for:

[0086] (16)

[0087] The jamming pulse and the radar pulse are set to be synchronized in time, and the amplitude of the jamming signal received by the radar is:

[0088] (17)

[0089] in, This indicates the interference power at the radar receiver. Indicates the jamming transmission power. Indicates the interfering carrier frequency. This indicates the gain of the jamming transmitting antenna, and the jammer's effect on the radar node. Interference probability for:

[0090] (18)

[0091] in, and They represent the first time. The interference airborne frequency is [frequency value] per pulse. Intermediate frequency bandwidth and radar nodes carrier frequency is The intermediate frequency bandwidth.

[0092] radar node In the The signal-to-interference-noise ratio (SINR) of each pulse is expressed as:

[0093] (19)

[0094] In this embodiment, step S2 is specifically as follows:

[0095] Step S21: Establish a generalized Markov decision process;

[0096] Each radar node in the radar system is considered an intelligent agent, and the radar cooperative anti-jamming process is modeled as a generalized Markov decision process, consisting of a quintuple. The specific definition is as follows:

[0097] (1) Intelligent agent set All intelligent frequency-agile radars constitute this intelligent entity set. ;

[0098] (2) Action set The selectable carrier frequencies of all radars constitute the action set. Any two radar nodes and of and They do not intersect; radar nodes In the Time step (i.e., the first) (pulse) action The launch of the The carrier frequency representation of each pulse, i.e. ;

[0099] (3) State set The state set consists of all selectable carrier frequencies of the jammer. Radar node In the state of time step Represented by the carrier frequency of the interference, i.e. ;

[0100] (4) State transition probability : Indicates radar node From state Execute action at time The state transitions to The transition probability is expressed as:

[0101] (20)

[0102] In this embodiment, It is considered unknown.

[0103] (5) Reward Set The reward set consists of the SINR of each pulse from all radars. Radar node In the The reward at each time step is represented as , It is obtained from equation (31).

[0104] Step S22: Obtain the state-action value function;

[0105] Set the optimal policy for each agent for:

[0106] (twenty one)

[0107] in, The state-action value function (also called the Q function) is defined as follows:

[0108] (twenty two)

[0109] in, This indicates the calculation of mathematical expectation. Indicates radar node In the Rewards at each time step Indicates the discount rate. Indicates the first After the first time step Each time step This represents the weighted value obtained from the discount rate. Indicates the first The discounted return at the time step is determined by future returns. We obtain the weighted sum.

[0110] like Figure 3 As shown in this embodiment, the PMAQL algorithm solution process in step S3 is as follows:

[0111] Step S31, let Initialize action set State set Discount rate Learning rate ;

[0112] Among them, learning rate Its value decreases as the iteration proceeds.

[0113] Step S32, Hypothesis: Co-training There are *coherent processing intervals* (CPIs), each CPI is considered an iteration, and in the *i*th... At the start of the next iteration, the initial actions are initialized randomly. Interference carrier frequency Obtain the initial state ;

[0114] Step S33: For each pulse in each CPI ,according to Greedy strategy selects action , ;

[0115] Step S34: Sensing the interfering carrier frequency from the environment. ;

[0116] Step S35, Integration and The reward was calculated. , get the state ;

[0117] Step S36: Update the intelligent agent of function ;

[0118] (twenty three)

[0119] in, They represent .

[0120] Step S37: Update Status ,when If the condition is met, return to step 33; otherwise, proceed to step S38.

[0121] Step S38, let ,like Otherwise, proceed to step S39;

[0122] Step S39: Output the final parallel output. Function, denoted as ;

[0123] (twenty four)

[0124] in, This indicates the action selection of multiple agents. Each time the radar pulse selects a carrier frequency, it searches for the detected interference carrier frequency. This allows us to obtain the optimal carrier frequency and achieve anti-interference.

[0125] The present invention also provides another embodiment, which verifies and analyzes the method of the present invention through simulation:

[0126] During target detection by a frequency-agile radar, the radar determines the carrier frequency of each pulse in the transmitted pulse train based on the carrier frequency of perceived environmental interference, and transmits a linear frequency modulation (LFM) signal. After being reflected by the target, the signal is received by the receiver, which feeds back whether interference has been detected in the pulse signal and performs signal processing to obtain the point information. Assuming there are three radar nodes in the radar system, and the frequency hopping ranges of the three radars are respectively... , and The instantaneous bandwidth of each radar is The number of pulses in one CPI transmitted by the radar is Each antenna has a transmit antenna gain and a receive antenna gain of 30dB, and a pulse width of [missing information]. PRT is Assume the target's initial position is directly above the radar at a distance of 160km, and... It flies horizontally at a speed of [speed value]. The target's RCS is [value]. The frequency hopping range of the sweeping jammer emitted by the target is... Instantaneous bandwidth is The frequency step value is Dry letter ratio Its frequency sweeping strategy involves sweeping frequencies in steps with a 70% probability and randomly selecting interference frequencies with a 30% probability. The reinforcement learning parameters are set as follows: the learning rate is... , Indicates the time step (i.e. the th time step) (number of pulses), discount rate The total number of training iterations was 100.

[0127] Figure 4 This section presents a comparison of the number of interference pulses under various methods. Since all radars exhibit similar performance, the results from radar 1 are used as representative. From... Figure 4 It can be seen that random frequency hopping and fixed carrier frequency methods perform poorly due to a lack of knowledge about unknown interference; in single-radar detection scenarios, approximately 50% of the pulses are jammed. Since the Q-learning-based frequency hopping method can learn the interference's frequency sweeping strategy and predict the next possible interference band based on the current carrier frequency, it achieves better anti-jamming performance than random and fixed carrier frequency methods. The proposed PMAQL algorithm has the best anti-jamming performance among all compared methods. While maintaining the same interference capability, it expands the frequency coverage of the radar signal; therefore, in a multi-radar system, the impact of interference on a single radar is smaller than in a Q-learning-based single-radar system.

[0128] Figure 5 The simulation results show a comparison of SINR across various methods. The simulation results demonstrate that the PMAQL algorithm achieves the highest signal-to-noise ratio. Specifically, the PMAQL algorithm improves SINR compared to all other methods, with an improvement of approximately 6 dB compared to Q-learning-based single-radar system methods. Figure 6 This represents the frequency bands of interference and radar in the last iteration of frequency resource scheduling using the PMAQL algorithm. Black indicates the interference band, gray indicates the radar pulse bands that are not interfered with, and light gray indicates the radar pulse bands that are interfered with. Figure 6 As can be seen from the PMAQL algorithm, each radar can use the current interference carrier frequency to predict the radar carrier frequency of the next time step, thereby effectively avoiding the interference frequency band.

[0129] In summary, the method of this invention can select a suitable frequency band for the radar based on the current frequency band of the interference, thereby avoiding the possible frequency band of the interference at the next moment. By modeling the multi-radar cooperative anti-interference process as a generalized Markov decision process, the radar carrier frequency is regarded as the action, the interference carrier frequency is regarded as the state, and SINR is used as the reward function. The radars cooperate with each other to reduce the interference degree of individual radars and effectively suppress frequency sweeping interference.

Claims

1. A frequency-agile multi-radar cooperative anti-jamming method based on reinforcement learning, the specific steps of which are as follows: Step S1: Establish the target echo model and interference signal model according to the phased array radar signal processing flow; Step S2: Treat each radar as an intelligent agent, model the radar cooperative anti-jamming process as a generalized Markov decision process, and obtain the state value function. Step S3: Solve the generalized Markov decision problem in step S2 using the parallel multi-agent Q-learning algorithm to obtain the radar carrier frequency selection strategy; The specific steps of S1 are as follows: The radar system is set to exist. Each radar node has a distance between any two radars that is negligible compared to the distance between the radar and the target. The transmitted signal from each radar node undergoes carrier frequency agility between pulses, and the number of pulses in a pulse train is... ; Indicates radar node The set of selectable carrier frequencies for the transmitted pulse. ; in, The first in Each element is represented as , This represents a fixed frequency step value, set relative to the radar bandwidth. Same size This indicates the number of selectable carrier frequencies for a radar pulse; and it is set that any two selectable carrier frequencies do not overlap, and the radar node... Transmitted pulse train signal Represented as: (1) in, Indicates time, Indicates the pulse repetition period. Indicates radar node The The carrier frequency of each pulse A linear frequency modulated signal representing a unit amplitude; Radar Node The transmitted signal illuminates the target, generating a target echo, which then triggers the radar node. In the The echo signal after being affected by interference and noise at each pulse point Represented as: (2) in, Indicates time delay. Indicates radar node With the goal in the Distance per pulse, Indicates the speed of propagation of electromagnetic waves. Indicates Doppler frequency shift, Let the power be zero mean. Gaussian white noise, Indicates unit interference signal, This indicates the change in amplitude caused by the propagation of the interference signal in space; Let represent a parameter that incorporates propagation effects and target scattering; its amplitude is expressed as: (3) in, Indicates radar node The The received power of each pulse at the radar receiver Indicates the transmission power. These represent the transmit antenna gain and the receive antenna gain, respectively. Indicates the signal wavelength. Represents the target's cross section (RCS); The jammer is configured to use a frequency sweeping strategy to interfere with radar signals. This indicates that the interference signal can be selected from a set of carrier frequencies; in, This indicates the number of selectable carrier frequencies for the interference signal. The first in Each element is represented as , This represents a fixed interference frequency step value, where the interference carrier frequency is randomly or according to a certain strategy from... Select the appropriate option to set the frequency hopping range of the interference to cover all possible frequency bands of the radar system; set the interference pulse and radar pulse to be synchronized in time, and the amplitude of the interference signal received by the radar to be: (4) in, This indicates the interference power at the radar receiver. Indicates the jamming transmission power. Indicates the interfering carrier frequency. This indicates the gain of the jamming transmitting antenna, and the jammer's effect on the radar node. Interference probability for: (5) in, and They represent the first time. The interference airborne frequency is [frequency value] per pulse. Intermediate frequency bandwidth and radar nodes carrier frequency is Intermediate frequency bandwidth; radar node In the The signal-to-interference-plus-noise ratio (SINR) of a pulse is expressed as: (6) Step S2 is as follows: Step S21: Establish a generalized Markov decision process; Each radar node in the radar system is considered an intelligent agent, and the radar cooperative anti-jamming process is modeled as a generalized Markov decision process, consisting of a quintuple. The specific definition is as follows: (1) Intelligent agent set All intelligent frequency-agile radars constitute this intelligent entity set. ; (2) Action set The selectable carrier frequencies of all radars constitute the action set. Any two radar nodes and of and They do not intersect; radar nodes In the Time step action The launch of the The carrier frequency representation of each pulse, i.e. ; (3) State set The state set consists of all selectable carrier frequencies of the jammer. Radar node In the state of time step Represented by the carrier frequency of the interference, i.e. ; (4) State transition probability : Indicates radar node From state Execute action at time The state transitions to The transition probability is expressed as: (7) in, Considered unknown; (5) Reward Set The reward set consists of the SINR of each pulse from all radars. Radar node In the The reward at each time step is represented as , It is obtained from equation (6); Step S22: Obtain the state-action value function; Set the optimal policy for each agent for: (8) in, The state-action value function, or Q-function, is defined as follows: (9) in, This indicates the calculation of mathematical expectation. Indicates radar node In the Rewards at each time step Indicates the discount rate. Indicates the first After the first time step Each time step This represents the weighted value obtained from the discount rate. Indicates the first The discounted return at the time step is determined by future returns. We obtain the weighted summation. In step S3, the specific solution process of the parallel multi-agent Q-learning algorithm is as follows: Step S31, let Initialize action set State set Discount rate Learning rate ; Among them, learning rate Its value decreases as the iteration proceeds; Step S32, Hypothesis: Co-training There are n coherent processing intervals CPI, each CPI is considered as one iteration, and in the nth iteration... At the start of the next iteration, the initial actions are initialized randomly. Interference carrier frequency Obtain the initial state , ; Step S33: For each pulse in each CPI , ,according to Greedy strategy selects action , ; Step S34: Sensing the interfering carrier frequency from the environment. ; Step S35, Integration and The reward was calculated. , get the state ; Step S36: Update the intelligent agent of function ; (10) in, They represent ; Step S37: Update Status ,when If the condition is met, return to step 33; otherwise, proceed to step S38. Step S38, let ,like Otherwise, proceed to step S39; Step S39: Output the final parallel output. Function, denoted as ; (11) in, This indicates the action selection status of multiple intelligent agents; each time the radar pulse selects a carrier frequency, it searches for the detected interference carrier frequency. This allows us to obtain the optimal carrier frequency and achieve anti-interference.

Citation Information

Patent Citations

  • Networking radar interference strategy generation method based on unmanned aerial vehicle group

    CN114911269A