A Design Method for Anti-interference Strategy of Frequency Modulation Agile Signal Based on DQN

By adopting a frequency-modulated agile signal anti-jamming strategy based on DQN, the problem of radar systems being unable to accurately detect targets in complex electromagnetic environments is solved. This enables the radar system to adaptively suppress interference while avoiding low signal-to-noise ratio and high probability of intercept caused by drastic changes in pulse width, thereby improving the radar's anti-jamming performance.

CN118465705BActive Publication Date: 2026-01-30UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410574401.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-10
Publication Date
2026-01-30
Estimated Expiration
2044-05-10

AI Technical Summary

Technical Problem

Radar systems struggle to accurately detect targets in complex electromagnetic environments, and existing anti-jamming technologies are unable to suppress interference while avoiding low signal-to-noise ratios and high probability of interception caused by drastic changes in pulse width.

Method used

A frequency-modulated agile signal anti-jamming strategy based on DQN is adopted. By modeling the interaction process between the radar and the jammer as a Markov decision process, a frequency-modulated agile signal model is designed. The optimal anti-jamming strategy of the radar is learned by using a deep Q-network. The modulation slope change of the radar signal is optimized by combining pulse width constraints and the reward function for jamming signal suppression.

Benefits of technology

This technology enables the radar system to adaptively suppress interference while avoiding low signal-to-noise ratio and high probability of interception caused by drastic changes in pulse width, thereby improving the radar's anti-jamming performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118465705B_ABST
    Figure CN118465705B_ABST
Patent Text Reader

Abstract

This invention discloses a design method for anti-jamming strategy of frequency-modulated agile signal based on DQN. First, the interaction process between the radar and the jammer is modeled as a Markov decision process. Then, the frequency modulation slope in the radar's transmitted waveform is used as the action, and the sequence of frequency modulation slopes in the radar and jammer's transmitted waveforms from the previous and current moments is used as the state. Next, a positive reward function is designed based on the correlation peak values ​​of the jamming and target signals, and a negative reward function is designed based on the signal energy and the probability of signal interception. Finally, the optimal SV-LFM signal slope variation strategy is learned through a deep Q-network. This method allows the radar to learn the optimal anti-jamming strategy without any prior information, enabling the radar system to adaptively suppress jamming while avoiding low signal-to-noise ratio and high interception probability caused by drastic pulse width changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of radar anti-jamming technology, specifically relating to a design method for frequency-modulated agile signal anti-jamming strategy based on DQN. Background Technology

[0002] With the development of electronic countermeasures technology, radar systems face complex electromagnetic environments and high threats, making it difficult for them to accurately detect targets. Therefore, developing adaptive radar anti-jamming technology is crucial to improving target detection performance. Reinforcement learning (RL), as a form of machine learning, is widely used in radar anti-jamming because it is a model-independent learning method that does not require prior knowledge of the environment. Current research can be broadly categorized into strategies for selecting anti-jamming measures and methods for designing anti-jamming waveform parameters.

[0003] Anti-jamming measure selection refers to the radar choosing the optimal anti-jamming measure for a specific jamming mode. To achieve autonomous selection of radar anti-jamming measures, the paper "W. Jiang, Y. Ren, and Y. Wang, 'Improving anti-jamming decision-making strategies for cognitive radar via multi-agent deep reinforcement learning,' DIGITAL SIGNAL PROCESSING, vol. 135, APR 30 2023" proposes a decision network based on the deep deterministic policy gradient (DDPG) algorithm. The authors explore the adversarial process between cognitive radar and intelligent jammers and demonstrate the effectiveness of multi-agent deep reinforcement learning (MDRL) for the adversarial decision-making system of cognitive radar.

[0004] Anti-jamming waveform parameter design refers to the process where radar, without selecting anti-jamming measures, directly alters the parameters of the transmitted waveform to achieve anti-jamming. Its core idea is to make the jamming parameters mismatched with the radar parameters, thereby preventing the jamming signal from accumulating gain at the radar receiver. Modifiable anti-jamming parameters include carrier frequency and pulse repetition frequency. The literature "K.Li, B.Jiu, and H.Liu, 'Deep q-network based anti-jamming strategy design for frequency agile radar,' in 2019 International Radar Conference (RADAR), 2019, pp.1-5" proposes a frequency-agile radar anti-jamming strategy design based on a deep Q-network. Its learning strategy enables the agent not only to avoid jamming but also to have a high detection probability. Many studies on anti-jamming waveform parameter design focus on carrier frequency. However, in addition to carrier frequency, anti-jamming can also be achieved by changing parameters such as frequency modulation slope, pulse repetition interval, and pulse width. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a design method for frequency-modulated agile signal anti-jamming strategy based on DQN, which enables the radar system to adaptively suppress interference while avoiding low signal-to-noise ratio and high interception probability caused by drastic pulse width changes.

[0006] The technical solution adopted in this invention is: a design method for anti-interference strategy of frequency modulation agile signal based on DQN, the specific steps of which are as follows:

[0007] Step 1: Establishing the frequency-modulated agile signal model;

[0008] If frequency-modulated agile signal SV-LFM is used as the radar signal transmission, then the expression for frequency-modulated agile signal s(t) is as follows:

[0009]

[0010] Where k+ξ represents the tuning slope of s(t), f0 represents the carrier frequency, t represents the fast time, and T represents the pulse repetition interval PRI; Let ξ represent a rectangular pulse function with a time width of t, k represent the fundamental frequency tuning slope, and ξ represent the random jitter parameter, i.e., the variation between the actual tuning slope and the fundamental frequency tuning frequency. Each transmitted pulse of the radar has a corresponding jitter parameter ξ, which is known only to the radar operator and not to the jamming party.

[0011] If a fixed signal bandwidth B ensures that the signal resolution remains constant, then the relationship between the signal pulse width and the frequency modulation slope is expressed as follows:

[0012]

[0013] Where τ represents the pulse width when the modulation frequency slope is k, and τ′ represents the pulse width when the modulation frequency slope is k+ξ. Under the same bandwidth, two SV-LFM signals with different slopes have different pulse widths.

[0014] The jamming signal is set as a deception signal coherent with the radar, intercepted, modulated, and forwarded by the jammer. The jammer forwards the currently intercepted radar signal each time. If the jitter parameters of the (m-1)th and mth radar pulses are ξ... m-1 and ξ m Then, when processing the m-th pulse, the matched filtering result S of the interference signal... o The expression for (t) is as follows:

[0015]

[0016] Step 2: Establish pulse width constraints;

[0017] By constraining the range of pulse width, it can be seen from the signal expression (1) that the average power P of each pulse is... R,av The expression is as follows:

[0018]

[0019] Where A represents the signal amplitude. The signal energy E is then expressed as follows:

[0020]

[0021] The input to the matched filter is set to a delayed signal plus additive white noise, as shown in the following expression:

[0022] x(t)=Cs(t-t1)+n(t) (6)

[0023] Where C represents a pre-set constant, t1 represents the time delay at the target distance, and n(t) represents the input white noise.

[0024] If the total noise power is set to N0, then the power spectral density is N0 / 2; using R h (t) represents the autocorrelation function of the filter, then the autocorrelation function of the output noise is... and power spectral density function They are represented as follows:

[0025]

[0026]

[0027] Where δ(t) represents the impulse response, and H(ω) represents the Fourier transform of the filter's impulse response h(t). The total average output noise power is equal to The value at t=0 is more precisely expressed as:

[0028]

[0029] The output signal power at time t is |Cs0(t-t1)| 2 And the output s of the signal is filtered. o From (t) = s(t) * h(t), we can obtain:

[0030]

[0031] The expression for the peak instantaneous signal-to-noise ratio at time 0 is as follows:

[0032]

[0033] Among them, E out This indicates the energy output by the radar receiver.

[0034] Let L be the signal energy attenuation during propagation. E,R Then E out =L E,R E.

[0035] When designing anti-jamming strategies, considering the probability of the jammer detecting the radar signal, the signal-to-noise ratio (SNR) output by the jammer receiver is... J,out The expression is as follows:

[0036]

[0037] Among them, P R,av L represents the average power of a single radar pulse transmission. R,J G represents the attenuation of pulse power from radar to jammer. J P represents the signal processing gain of the jammer receiver. J,n This represents the power of the jammer's receiver noise. The detection probability of the jammer is approximated as:

[0038]

[0039] Among them, P fa represents the false alarm probability of the jammer, and erfc represents the complementary error function.

[0040] Step 3: Based on the established frequency modulation agile signal model, establish an anti-interference scenario model based on MDP;

[0041] The adversarial process between radar and jammer is modeled as an MDP model, and the radar anti-jamming strategy problem is expressed as an MDP problem.

[0042] At each time step t, the radar selects the modulation slope of the transmitted SV-LFM signal according to a given strategy π. Simultaneously, as part of the environment, the jammer relays the signal from the previous pulse transmission cycle, with a modulation slope of a. t j Subsequently, the state changed from s t Transfer to s t+1 The agent will receive a reward based on its state and behavior. t+1 .

[0043] Step 4: Based on the MDP model in Step 3, set the action, state, and reward functions, construct the DQN model, train and solve it to generate the optimal anti-interference strategy;

[0044] Based on the MDP problem in step three, DQN is used to solve it.

[0045] Using a function Q(s,a) to approximate the action value Q(s,a;ω) is called value function approximation. Then, a neural network is used to generate the function Q(s,a;ω), which is called DQN. ω represents the training parameters of the neural network.

[0046] The action, state, and reward function settings for the radar and jammer MDP problem are as follows:

[0047] The SV-LFM signal with a slope selection range of [-50%, 50%] is divided into n = 50 discrete points as its action in each time slice, while the jammer's signal slope is... That is, the modulation slope of a single pulse on the radar.

[0048] Then select As a state sequence, this state sequence includes historical information from the radar and jammer.

[0049] In the reward setting, the reward function is divided into a positive reward obtained by suppressing interference signals through matched filtering and a negative reward obtained by pulse width constraint. After matched filtering of the echo signal, the interference signal will be orthogonal to the current radar signal. Therefore, the absolute value of the cross-correlation normalized amplitude is taken as the positive reward, i.e., r. t 1 =|S o (t)|.

[0050] The pulse width is limited; when the pulse width is too narrow, the energy of the transmitted pulse signal is selected as a negative reward. Furthermore, the change in pulse width linearly alters the energy value of the sub-signal; therefore, r is set... t 2 =(1 / E) 2This negative reward reflects the constraint on narrow pulses rather than wide pulses. When the pulse width is too large, the probability of the jammer intercepting the radar signal is chosen as the negative reward, i.e., r. t 3 =P d .

[0051] Then define a reward function, denoted as r. t =r t 1 -ar t 2 -br t 3 The coefficients a and b are varied according to the requirements of the pulse width variation range.

[0052] Using the actions, states, and reward functions set above, the MDP problem in step three is solved through iterative DQN.

[0053] The specific steps for each iteration of the DQN method to solve the MDP problem are as follows:

[0054] (1) Select action a using the ∈-greedy strategy t Observe the disruptor's strategy

[0055] (2) Use a t and Calculate reward r t ;

[0056] (3) Obtain the next state s t+1 ;

[0057] (4) Store the array (s) in the experience replay array D. t ,a t ,r t ,s t+1 );

[0058] (5) Update ω = ω + ▽ using gradient gradation. ω (y t -Q(s t ,a t ∣ω)) 2 ;

[0059] Among them, ▽ ω This indicates that the gradient with respect to ω is calculated. ω - The parameter represents the target network parameter, where parameter γ (0≤γ≤1) represents the discount factor.

[0060] (6) Pre-set the update step number L, and update ω every L steps. - For ω;

[0061] (7) Repeat steps (1)-(6) until the total reward value converges, and obtain the optimal anti-interference strategy through iterative iteration.

[0062] The beneficial effects of this invention are as follows: First, the method of this invention models the interaction process between the radar and the jammer as a Markov decision process. Then, it uses the frequency modulation slope in the radar's transmitted waveform as the action, and the sequence of frequency modulation slopes in the radar and jammer's transmitted waveforms from the previous and current moments as the state. Next, a positive reward function is designed based on the correlation peak values ​​of the jamming and target signals, and a negative reward function is designed based on the signal energy and the probability of signal interception. Finally, a deep Q-network is used to learn the optimal SV-LFM signal slope variation strategy. This method allows the radar to learn the optimal anti-jamming strategy without any prior information, enabling the radar system to adaptively suppress jamming while avoiding low signal-to-noise ratio and high interception probability caused by drastic pulse width changes. Attached Figure Description

[0063] Figure 1 This is a flowchart of a frequency-modulated agile signal anti-interference strategy design method based on DQN according to the present invention.

[0064] Figure 2 This is an example of SV-LFM autocorrelation and cross-correlation plots in an embodiment of the present invention.

[0065] Figure 3 This is a diagram illustrating the interaction process between the radar and the jammer in an embodiment of the present invention.

[0066] Figure 4 This is a block diagram of the DQN algorithm based on the target network and experience replay in an embodiment of the present invention.

[0067] Figure 5 This is a graph showing the total reward convergence curves for the optimal and random strategies in this embodiment of the invention.

[0068] Figure 6 This is a slope point map generated by the optimal strategy in this embodiment of the invention.

[0069] Figure 7 This is a slope point plot generated by a random strategy in an embodiment of the present invention.

[0070] Figure 8 This is the pulse width graph generated by the optimal strategy in this embodiment of the invention.

[0071] Figure 9 This is a diagram showing the matched filtering results of some pulses under random and optimal strategies in an embodiment of the present invention. Detailed implementation method:

[0072] To facilitate the description of the method of this invention, the following terms are explained.

[0073] Term 1: Markov decision process (MDP);

[0074] Mathematical Discrete Decision Processes (MDPs) are a mathematical framework for describing stochastic decision problems. MDPs are primarily used to describe decision processes characterized by stochasticity, uncertainty, and multi-step decision-making. In an MDP, the decision problem is modeled as including a state *s*, an action *a*, a reward *r*, and transit probabilities. The mathematical framework of MDP. The state transitions in MDP satisfy the Markov property, that is, the future state is only related to the current state and is independent of the past state.

[0075] Term 2: Reinforcement Learning (RL);

[0076] Reinforcement learning is a method for handling dynamic decision-making problems, where an agent interacts with its environment and makes different decisions to achieve an optimal or specific goal. Value function algorithms are one of the important algorithms in reinforcement learning, with Q-learning being a typical example.

[0077] The effectiveness of the method of this invention is verified through simulation experiments. All steps and results of the method are verified on the PyCharm simulation platform. The method of this invention will be further described below with reference to the accompanying drawings and embodiments.

[0078] like Figure 1 The flowchart shown below illustrates a design method for a frequency-modulated agile signal anti-interference strategy based on DQN according to the present invention. The specific steps are as follows:

[0079] Step 1: Establishing the frequency-modulated agile signal model;

[0080] Considering that radar systems transmit conventional pulse linear frequency modulated (LFM) signals, DRFM jammers can intercept, sample, and store these signals. Therefore, DRFM jammers can reconstruct copies of the transmitted radar signals, creating range decoys to confuse the radar system. To counter DRFM repetition jamming, slope-varying linear frequency modulation (SV-LFM) signals are used for radar signal transmission.

[0081] In this embodiment, the radar system transmits a frequency-modulated agile signal, the expression of which is as follows:

[0082]

[0083] Where k+ξ represents the tuning slope of s(t), f0 represents the carrier frequency, t represents the fast time, and T represents the pulse repetition interval PRI; This represents a rectangular pulse function with a time width of t. It's important to emphasize that k represents the fundamental frequency tuning slope, and ξ represents the random jitter parameter, i.e., the variation between the actual tuning slope and the fundamental frequency tuning frequency. Each transmitted pulse from the radar has a corresponding jitter parameter ξ, which is known only to the radar operator and not to the jamming party.

[0084] In this embodiment, to ensure that the signal resolution remains constant, the signal bandwidth B needs to be fixed. Therefore, the relationship between the signal pulse width and the frequency modulation slope is expressed as follows:

[0085]

[0086] Where τ represents the pulse width when the modulation frequency slope is k, and τ′ represents the pulse width when the modulation frequency slope is k+ξ. Under the same bandwidth, two SV-LFM signals with different slopes have different pulse widths.

[0087] The jamming signal is set as a deception signal coherent with the radar, intercepted, modulated, and forwarded by the jammer. The jammer forwards the currently intercepted radar signal each time. If the jitter parameters of the (m-1)th and mth radar pulses are ξ... m-1 and ξ m Then, when processing the m-th pulse, the matched filtering result S of the interference signal... o The expression for (t) is as follows:

[0088]

[0089] As can be seen from equation (3), the amplitude of the interference signal is related to the jitter difference of the modulation frequency slopes of the two signals. The greater the jitter difference of the tuning frequency, the smaller the output amplitude of the matched filter of the interference signal.

[0090] like Figure 2 As shown, the correlation between different frequency modulation jitter signals and the original signal is illustrated. It can be seen that the larger the slope jitter, the smaller the correlation between the SV-LFM signal and the matched signal, that is, the smaller the output of the matched filter.

[0091] Step 2: Establish pulse width constraints;

[0092] Since the signal bandwidth considered in step one is constant, the pulse width of the signal will vary with the slope. An excessively narrow pulse width leads to a low signal-to-noise ratio, while an excessively large pulse width increases the probability of radar signal interception. Therefore, it is necessary to constrain the range of pulse width to ensure the radar has good anti-jamming capabilities while maintaining performance. The impact of pulse width variation is analyzed below from the perspective of transmitted signal energy.

[0093] From the signal expression (1), it can be seen that the average power P of each pulse is... R,av The expression is as follows:

[0094]

[0095] Where A represents the signal amplitude. The signal energy E is then expressed as follows:

[0096]

[0097] The input to the matched filter is set to a delayed signal plus additive white noise, as shown in the following expression:

[0098] x(t)=Cs(t-t1)+n(t) (6)

[0099] Where C represents a pre-set constant, t1 represents the time delay at the target distance, and n(t) represents the input white noise.

[0100] If the total noise power is set to N0, then the power spectral density is N0 / 2; using R h (t) represents the autocorrelation function of the filter, then the autocorrelation function of the output noise is... and power spectral density function They are represented as follows:

[0101]

[0102]

[0103] Where δ(t) represents the impulse response, and H(ω) represents the Fourier transform of the filter's impulse response h(t). The total average output noise power is equal to The value at t=0 is more precisely expressed as:

[0104]

[0105] The output signal power at time t is |Cs0(t-t1)| 2 And the output s of the signal is filtered. o From (t) = s(t) * h(t), we can obtain:

[0106]

[0107] The expression for the peak instantaneous signal-to-noise ratio at time 0 is as follows:

[0108]

[0109] Among them, E out This indicates the energy output by the radar receiver.

[0110] In this embodiment, the signal energy attenuation during propagation is set to L. E,R Then E out =L E,RE. As can be seen from equations (9) and (11), when the pulse width decreases, the energy of the transmitted signal also decreases, resulting in a decrease in the signal-to-noise ratio of the output signal after matched filtering. Therefore, it is important to limit the pulse width to a relatively narrow range, that is, the frequency modulation slope of SV-LFM should not be too large.

[0111] Because the jammer intercepts radar signals and performs carrier parameter identification during the attack, it will only forward the jamming pulse in subsequent transmissions after successfully intercepting the radar transmitter signal. Therefore, the probability of the jammer detecting the radar signal must be fully considered when designing anti-jamming strategies. The jammer receiver output signal-to-noise ratio (SNR) is also important. J,out The expression is as follows:

[0112]

[0113] Among them, P R,av L represents the average power of a single radar pulse transmission. R,J G represents the attenuation of pulse power from radar to jammer. J P represents the signal processing gain of the jammer receiver. J,n This represents the power of the jammer's receiver noise. The detection probability of the jammer is approximated as:

[0114]

[0115] Among them, P fa The signal represents the false alarm probability of the jammer, and ERFC represents the complementary error function. When the pulse width is too large, the probability of the jammer detecting our radar signal also increases, which will increase the probability of the jammer being intercepted. Therefore, the pulse width cannot be too wide, that is, the frequency modulation slope of the SV-LFM signal cannot be too small.

[0116] Step 3: Based on the established frequency modulation agile signal model, establish an anti-interference scenario model based on MDP;

[0117] This embodiment models the confrontation process between the radar and the jammer as an MDP model.

[0118] At each moment, the agent selects an action based on the current state, and after performing the action, it moves to the next state and receives a reward, while the environment changes to another state with a certain probability.

[0119] In MDP, a reward function is introduced to measure the instantaneous reward obtained by the agent after performing an action in each state. The policy in MDP describes which actions the agent should choose in each state.

[0120] like Figure 3As shown, the radar anti-jamming strategy problem in this embodiment can be formulated as an MDP problem. At each time step t, the radar selects the modulation slope of the transmitted SV-LFM signal according to a given strategy π. Simultaneously, as part of the environment, the jammer relays the signal from the previous pulse transmission cycle, whose modulation slope is... Subsequently, the state changed from s t Transfer to s t+1 The agent will receive a reward based on its state and behavior. t+1 .

[0121] The goal of this invention is to design an appropriate modulation slope for the SV-LFM signal to maximize the long-term cumulative reward.

[0122] Step 4: Based on the MDP model in Step 3, set the action, state, and reward functions, construct the DQN model, train and solve it to generate the optimal anti-interference strategy;

[0123] Based on the MDP problem in step three, DQN (deep Q-network) is used to solve it.

[0124] Q-learning is a table-based policy optimization algorithm. Its table represents the correspondence between states and actions, with each entry representing the value of the action taken in that state. The learning process of Q-learning first initializes the table by setting all Q(s,a) to 0. Then, it selects initial states and actions according to a greedy policy and iteratively updates the value function in the table using the Bellman equation until the algorithm converges. After training, the algorithm makes decisions based on the action with the highest Q-function value, gradually approaching the final goal. The iterative formula of Q-learning is expressed as:

[0125] Q(s,a)←Q(s,a)+α[R+γmax a′ Q(s′,a′)-Q(s,a)](14)

[0126] The iterative formula for Q-learning is called the Bellman equation, which describes how the value function of the current state is updated according to the value function of the next state. The left side, Q(s,a), represents the new Q-value of action a in state s, calculated in one iteration. The right side, Q(s,a), represents the original Q-value, calculated in the last iteration. R represents the reward obtained after performing action a, indicating the feedback received from performing action a in state s. a′Q(s′,a′) represents the maximum Q-value among all possible actions a′ in the next state s, indicating the next Q-value of the optimal action the agent can choose in the next state s′. The parameter α (0≤α≤1) represents the learning rate, which controls the weight between new and old values; in simplified form, α = 1. The parameter γ (0≤γ≤1) represents the discount factor, indicating the importance of future rewards in calculating the maximum Q-value of the next state s′. Q-learning iteratively updates the value function, gradually optimizing the policy to enable the agent to better handle dynamic decision-making problems.

[0127] The Q-learning algorithm maintains a Q-table to store the reward obtained by taking action a in each state s, i.e., the state-value function Q(s,a). However, it has significant limitations. In many real-world scenarios, reinforcement learning tasks involve continuous state spaces with an infinite number of states. In such cases, the value function can no longer be stored using a table. To address this issue, this invention approximates the action value Q(s,a; ω) using a function Q(s,a), known as the value function approximation. Then, a neural network is used to generate the function Q(s,a; ω), called Deep Q-network (DQN), where ω represents the neural network training parameters.

[0128] The DQN flowchart is as follows: Figure 4 As shown. The experience replay array D can decompose the sequence, eliminate correlations, and ensure the data follows an independent uniform distribution. This reduces the variance of parameter updates and improves convergence speed. It also allows reuse of experiences with high data utilization. When the weights change, the estimated action values ​​also change; the target network is designed to mitigate instability issues. ω - This represents the parameters of the target network.

[0129] The action, state, and reward function settings for the radar and jammer MDP problem are as follows:

[0130] Since DQN cannot handle continuous variables, this embodiment divides the SV-LFM signal with a slope selection range of [-50%, 50%] into 50 discrete points as its action in each time slice, while the jammer's signal slope is... This refers to the modulation slope of a single pulse on the radar. This embodiment selects... As a state sequence, this state sequence contains historical information from the radar and jammer, which helps the agent learn the optimal policy.

[0131] In the reward setting, this embodiment divides the reward function into a positive reward obtained by suppressing interference signals through matched filtering and a negative reward obtained by pulse width constraint. Since the interference signal will be orthogonal to the current radar signal after matched filtering of the echo signal, this embodiment takes the absolute value of the cross-correlation normalized amplitude as the positive reward, i.e., r t 1 =|S0(t)|.

[0132] The pulse width is limited; when the pulse width is too narrow, this embodiment selects the energy of the transmitted pulse signal as a negative reward. However, since the change in pulse width linearly alters the signal energy value, this embodiment sets r... t 2 =(1 / E) 2 This reflects the constraint that the negative reward applies to narrow pulses rather than wide pulses. In this embodiment, when the pulse width is too large, the probability of the jammer intercepting the radar signal is selected as the negative reward, i.e., r. t 3 =P d .

[0133] Therefore, this embodiment defines a reward function, denoted as r. t =r t 1 -ar t 2 -br t 3 The coefficients a and b can be varied according to the required range of pulse width variation.

[0134] Using the actions, states, and reward functions set above, the MDP problem in step three is solved through iterative DQN.

[0135] The specific steps for each iteration of the DQN method to solve the MDP problem are as follows:

[0136] (1) Select action a using the ∈-greedy strategy t Observe the disruptor's strategy

[0137] (2) Use a t and Calculate reward r t ;

[0138] (3) Obtain the next state s t+1 ;

[0139] (4) Store the array (s) in the experience replay array D. t ,a t ,r t ,s t+1 );

[0140] (5) Update ω = ω + ▽ using gradient gradation. ω (y t -Q(s t ,a t ∣ω)) 2 ;

[0141] Among them, ▽ ω This indicates that the gradient with respect to ω is calculated. ω - The parameter represents the target network parameter, where parameter γ (0≤γ≤1) represents the discount factor.

[0142] (6) Pre-set the update step number L, and update ω every L steps. - For ω;

[0143] (7) Repeat steps (1)-(6) until the total reward value converges, and obtain the optimal anti-interference strategy through iterative iteration.

[0144] In this embodiment, to verify the effectiveness and fairness of the method of the present invention, each iteration is performed 200 times, for a total of 400 iterations, and 100 Monte Carlo experiments are conducted. The experimental parameters are shown in Table 1.

[0145] Table 1

[0146]

[0147] Figure 5 The convergence curves of the total reward for the optimal policy and the randomized policy are presented. The total reward is determined by the sum of the settling time (rt) over each 200 iterations. The shaded area in the figure represents the results of multiple Monte Carlo simulations, and the solid line represents the average of multiple experiments. It can be seen that the total reward of the optimal policy increases with the number of iterations, eventually converging after 200 iterations. The total reward is approximately 3000. In contrast, the randomized policy does not adaptively change its policy with the number of iterations, and the total reward remains consistently around 1200. Therefore, the performance of the optimal policy is significantly higher than that of the randomized policy.

[0148] Figure 6 and Figure 7 The images show the slope points of 10 pulses for both the optimal and randomized strategies. It can be seen that the optimal strategy exhibits a frequency modulation slope that jumps continuously between two slope points. However, the randomized strategy shows a wider range of slope point jumps. Furthermore, the pulse width varies with the slope. Figure 8 As shown, the pulse width of the randomized strategy fluctuates significantly. However, the pulse width of the optimal strategy jumps between two points, and its range of variation is smaller than that of the randomized strategy.

[0149] Figure 9The matched filtering results for some pulses under random and optimal strategies are shown. Comparing the anti-jamming capabilities of the three pulses, it is clear that the 8th pulse of the randomized strategy has worse anti-jamming capability than the optimal strategy. The 2nd pulse of the randomized strategy is slightly better than the optimal strategy, but the pulse width of the radar signal is too narrow. The received signal-to-noise ratio is low, which is not conducive to target detection. Since our method balances anti-jamming capability and pulse width constraints, its overall performance is better than that of the randomized strategy.

[0150] In summary, the method of this invention mainly utilizes deep Q-networks in reinforcement learning to design a reward function that combines positive rewards for suppressing interference with negative rewards for suppressing pulse width. This solves the problem of neglecting the reduction in signal-to-noise ratio and the increase in the probability of interception in order to achieve anti-interference capability. The radar-jamming confrontation scenario is modeled as a Markov decision process. The positive reward function is designed based on the correlation peak values ​​between the jamming signal and the target signal, while the negative reward function is designed based on the signal energy and the probability of signal interception. Based on the obtained anti-interference strategy, the radar system can adaptively suppress interference while avoiding low signal-to-noise ratio and high interception probability caused by drastic pulse width changes.

[0151] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. A DQN-based frequency agile signal anti-jamming strategy design method, the specific steps are as follows: Step one, frequency agile signal model establishment; When the frequency-modulated signal SV-LFM is used as the radar signal transmission, the frequency-modulated signal The expression is as follows: (1); wherein express The tuning slope, Indicates the carrier frequency. Indicates a fast time. Indicates the pulse repetition interval PRI; Indicates time width as rectangular pulse function, This represents the fundamental frequency tuning slope. This represents the random jitter parameter, i.e., the variation between the actual frequency modulation slope and the fundamental tuning frequency; each transmitted pulse of the radar has a corresponding jitter parameter. This parameter is known only to the radar operator, not to the jammer. Fixed signal bandwidth If the resolution of the signal is guaranteed to be constant, the relationship between the signal pulse width and the frequency modulation slope is expressed as follows: (2); wherein, denotes the pulse width when the modulation frequency slope is denotes the pulse width when the modulation frequency slope is ; under the same bandwidth, the SV-LFM signals with two different slopes have different pulse widths;​ The jamming signal is set to a deception signal coherent with the radar, which is intercepted, modulated, and forwarded by the jammer; the jammer forwards the currently intercepted radar signal each time; if the radar... The, the The jitter parameters of each pulse are as follows: and Then in processing the first Matched filtering result of interference signal during pulse. The expression is as follows: (3); Step two, pulse width constraint condition establishment; The range of the pulse width is restricted, and from the signal expression (1), the average power of each pulse The expression is as follows: (4); wherein represents the signal amplitude; the energy of the signal The expression is as follows: (5); Set the input of the matched filter as a delayed signal plus additive white noise, the expression is as follows: (6); wherein, represents a constant set in advance, represents a time delay at a target distance, represents an input white noise; The total noise power is set as The power spectral density is ; the autocorrelation function of the filter is represented as The autocorrelation function of the output noise and the power spectral density function are represented as (7); (8); wherein denotes the impulse response, denotes the Fourier transform of the filter impulse response ; the total average output noise power is equal to the value at is more precisely expressed as: (9); At the output signal power at the moment is and the filter output of the signal is It can be obtained: (10); Then the peak instantaneous signal-to-noise ratio expression of the output at the 0th moment is as follows: (11); wherein represents the energy output by the radar receiver; The attenuation of the set signal energy during propagation is then ; In designing the anti-jamming strategy, considering the probability that the jammer detects the radar signal, the output signal-to-noise ratio of the jammer receiver is The expression is as follows: (12); wherein, Ppresents the average power of a single radar pulse transmission, Ppresents the attenuation of the pulse power from the radar to the jammer, Ppresents the signal processing gain of the jammer receiver, Ppresents the power of the jammer receiver noise; the detection probability of the jammer is approximated as: (13); wherein, denotes the false alarm probability of the jammer, denotes the complementary error function; Step three, according to the established frequency agile signal model, establish the anti-jamming scene model based on MDP; Model the confrontation process between the radar and the jammer as an MDP model, and express the radar anti-jamming strategy problem as an MDP problem; At each time step , the radar selects the modulation slope of the transmitted SV-LFM signal according to the given strategy ; meanwhile, as part of the environment, the jammer retransmits the signal of the previous pulse transmission period, whose modulation slope is ; subsequently, the state is transferred from to , and the agent will obtain a reward according to the state and action. Step four, based on the MDP model of step three, set the action, state and reward function, build the DQN model for training and solving, and generate the optimal anti-jamming strategy; Based on the MDP problem of step three, DQN is used to solve it; The function is used to approximate the action value , called value function approximation, and then a neural network is used to generate the function , called DQN, denotes the neural network training parameters; The action, state and reward function settings in the radar and jammer MDP problem are as follows: The SV-LFM signal with a slope selection range of is divided into discrete points as its action in each time slice, while the signal slope of the jammer is , that is, the modulation slope of the previous pulse on the radar; Then select As the state sequence, the state sequence includes historical information of the radar and the jammer; In the reward setting, the reward function is divided into positive reward obtained by suppressing the interference signal through matched filtering and negative reward obtained by pulse width constraint; after matched filtering of the echo signal, the interference signal will be orthogonal to the current radar signal, so the absolute value of the normalized amplitude of mutual correlation is taken as the positive reward, that is ​ The pulse width is limited, when the pulse width is too narrow, the energy of the selected transmitting pulse signal is selected as a negative reward; and the change of the pulse width linearly changes the energy value of the signal, so that Reflecting this negative reward is a constraint for narrow pulses rather than wide pulses; when the pulse width is too large, the probability of the jammer intercepting the radar signal is selected as a negative reward, that is ; Then define a reward function, expressed as ; coefficient , According to the requirement of pulse width variation range Through the above setting of the action, state and reward function, the MDP problem in step three is solved by using the cycle iteration of DQN; Then the specific steps of solving the MDP problem by using DQN in each cycle are as follows: (1) select an action with the e-greedy policy , observe the interferer policy ; (2) Use and calculate rewards ; (3) Obtain next state ; (4) storing the array in an experience replay array ;​ (5) Gradient update ; wherein, denotes the gradient of , , denotes the parameters of the target network, the parameters denotes a discount factor, ; (6) Pre-set the number of steps L, update every L steps For ; (7) Loop steps (1)-(6) until the total reward value converges, and the optimal anti-jamming strategy is obtained through cycle iteration.