A joint time-frequency domain jamming method for frequency-agile radar based on fuzzy reinforcement learning

By using fuzzy reinforcement learning combined with a time-frequency joint jamming algorithm, the jamming duration and frequency selection are optimized, solving the problem of insufficient anti-jamming performance of traditional jamming strategies for frequency-agile radar in the face of intelligent radar, and achieving a more efficient jamming effect.

CN118311508BActive Publication Date: 2025-12-02UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410415598.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-08
Publication Date
2025-12-02
Estimated Expiration
2044-04-08

AI Technical Summary

Technical Problem

When facing intelligent radars, existing frequency-agile radars are unable to achieve effective joint jamming in the time and frequency domains using traditional jamming strategies, resulting in insufficient anti-jamming performance.

Method used

A fuzzy reinforcement learning-based approach is adopted to model the jammer as an intelligent agent. Through a generalized Markov decision process, combined with time-domain fuzzy Q-learning and frequency-domain Q-learning algorithms, the jamming duration and frequency selection are optimized to achieve joint time-frequency domain jamming.

Benefits of technology

It improves the jammer's anti-jamming capability against frequency-agile radar, and can select appropriate jamming frequency bands and durations according to the radar frequency band and pulse repetition period, effectively avoiding the frequency band at the next moment and improving the jamming effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118311508B_ABST
    Figure CN118311508B_ABST
Patent Text Reader

Abstract

This invention discloses a time-frequency domain joint jamming method for frequency-agile radar based on fuzzy reinforcement learning. First, according to the radar signal processing flow, a target echo model and a jamming signal model are established. The jammer is treated as an intelligent agent, and the time-frequency domain joint jamming process is modeled as a generalized Markov decision process, yielding a state-value function. This function is then solved using a proposed time-domain fuzzy Q-learning algorithm and a frequency-domain Q-learning algorithm, finally obtaining the jammer's jamming time and carrier frequency selection strategy. The method of this invention models the jamming process as a generalized Markov decision process. In the frequency domain, the jammer's carrier frequency selection is considered an action, and the radar carrier frequency is considered a state. In the time domain, the jammer's jamming duration selection is considered an action, and each fuzzy rule is considered a state. The product of the correlation entropy between the time-domain inferred jamming duration and the radar pulse repetition time, and the frequency-domain jamming JNSR, is used as the joint reward function to achieve a time-frequency domain joint jamming effect, effectively implementing jamming.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of radar jamming technology, specifically relating to a time-frequency domain joint jamming method for frequency-agile radar based on fuzzy reinforcement learning. Background Technology

[0002] Frequency-agile radar, as a typical electronic countermeasures device, has been widely used in various electronic warfare scenarios. It achieves anti-jamming by rapidly changing the signal carrier frequency within, between, or between pulse groups, and features long detection range, high angle measurement accuracy, and strong anti-jamming capability. Frequency sweeping jamming, as a commonly used electronic countermeasures technique, reduces radar resolution and detection performance by dynamically scanning radar frequencies. With the enhancement of jamming capabilities and the development of cognitive jamming, the random or pseudo-random frequency hopping strategies of frequency-agile radar can no longer achieve effective anti-jamming performance. Therefore, researching a new and effective anti-jamming technology is an urgent need for future electronic warfare applications.

[0003] With the rapid development of radar intelligence, many radar systems have acquired autonomous learning capabilities, and the combat effectiveness of radar countermeasures against conventional radars is declining. Therefore, it is necessary to consider the intelligentization of jamming methods. The literature "Zhu Bakun, Zhu Weigang, Li Wei, et al. A review of radar jamming decision-making technology based on reinforcement learning. Electro-Optics and Control, 2022, 29(04): 52-58+111" reviews traditional radar jamming decision-making algorithms, elaborates on the principle of using reinforcement learning for radar jamming decision-making, and analyzes the current development status of radar jamming decision-making technology based on reinforcement learning.

[0004] The paper "X. Qiang, Z. Weigang and J. Xin, 'Research on method of intelligent radar confrontation based on reinforcement learning', 2017 2nd IEEE International Conference on Computational Intelligence and Applications (ICCIA), Beijing, China, 2017, pp. 471-475" describes how, to counter intelligent radar, jammers generate intelligent jamming strategies through interaction with the external environment. The jamming process is modeled as a Markov Decision Process (MDP), and a reinforcement learning algorithm is used to solve for the optimal strategy. This approach shows better performance compared to existing radar countermeasures. However, considering only effective jamming in the frequency domain, current research on multi-domain joint jamming algorithms based on reinforcement learning is limited; therefore, further research is necessary. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention proposes a time-frequency domain joint jamming method for frequency-agile radar based on fuzzy reinforcement learning.

[0006] The technical solution adopted in this invention is: a time-frequency domain joint jamming method for frequency-agile radar based on fuzzy reinforcement learning, the specific steps of which are as follows:

[0007] S1. Based on the radar signal processing flow, establish the target echo model and the interference signal model;

[0008] S2. Treat the jammer as an intelligent agent and model the jamming process of the jammer jamming the frequency-agile radar as a generalized Markov decision process, and obtain the state value function of frequency domain Q-learning and the state value function of time domain fuzzy Q-learning respectively.

[0009] S3. Solve the generalized Markov decision problem in step S2 using the agent time-domain fuzzy Q-learning algorithm and the frequency-domain Q-learning algorithm to obtain the jammer's jamming duration selection strategy and jammer frequency domain selection strategy, ensuring that the jammer can achieve joint jamming in the time and frequency domains after the radar frequency changes.

[0010] Furthermore, step S1 is specifically as follows:

[0011] The radar's transmitted signal is configured to undergo carrier frequency agility between pulses, with the number of pulses in a pulse train being L. N represents the set of selectable carrier frequencies for radar transmit pulses. R This indicates the total number of carrier frequencies that can be selected for radar transmission pulses, and sets that any two selectable carrier frequencies do not overlap.

[0012] in, The i-th element in is represented as f i =f i-1 +Δf R ,i∈{2,3,…,N R}, f i f represents the selection of the frequency value of the i-th element; i-1 Indicates the selection of the frequency value of the (i-1)th element; Δf R This represents a fixed frequency step value, with the radar bandwidth set to B. R The bandwidth for suppressing interference signals is B. J The relationship between the two is B. J =5~10B R The expression for the radar's transmitted pulse train signal x(t) is as follows:

[0013]

[0014] Where t represents time T rf represents the pulse repetition period. n Let represent the carrier frequency of the nth pulse of the radar, and s(t) represent the linear frequency modulated signal with unit amplitude.

[0015] When a radar transmits a signal to illuminate a target, generating a target echo, the echo signal y after the radar is affected by interference and noise at the nth pulse will be considered as follows. n The expression for (t) is as follows:

[0016]

[0017] Where, τ n =2R n / c indicates delay, R n Let f represent the radial distance between the radar and the target at the nth pulse, c represent the propagation speed of the electromagnetic wave, and f represent the radial distance between the radar and the target at the nth pulse. d,n N represents the Doppler frequency shift. n (t) represents a zero-mean power P. N Gaussian white noise, J n (t) represents the unit interference signal, β n α represents the change in amplitude caused by the propagation of the interference signal in space. n Let represent a parameter that incorporates propagation effects and target scattering; its amplitude expression is as follows:

[0018]

[0019] Among them, P R,n P represents the received power of the nth pulse at the radar receiver. T G represents the radar transmit power. T G R λ and σ represent the radar transmitting antenna gain and receiving antenna gain, respectively, λ represents the signal wavelength, and σ represents the target cross section (RCS).

[0020] The jammer is configured as an agent that generates a jamming duration strategy in the time domain through fuzzy reinforcement learning. The measured pulse repetition period (PRT) is input into the fuzzy inference system to produce a pulse repetition period output value with smaller error, which is the jammer's jamming duration and serves as the global action selection in the time domain.

[0021] The jammer is configured as an intelligent agent that generates a strategy in the frequency domain through reinforcement learning to jam radar signals. This indicates that the interference signal can be selected from a set of carrier frequencies.

[0022] Where, N J This indicates the number of selectable carrier frequencies for the interference signal. The p-th element in is represented as f J,p =f J,p-1 +ΔfJ , Δf J This represents a fixed interference frequency step value, where the interference carrier frequency is randomly or according to a certain strategy from... Select the appropriate option to set the frequency hopping range of the interference to cover all possible frequency bands of the radar system; the amplitude expression of the interference signal received by the radar is as follows:

[0023]

[0024] Among them, P I,n P represents the interference power at the radar receiver. J f represents the jamming transmission power. J,n Indicates the interfering carrier frequency, G J G represents the gain of the interfering transmitting antenna. R U represents the radar receiving antenna gain, and the probability of interference from the jammer to the radar. n The expression is as follows:

[0025]

[0026] Among them, B J,n (f J,n ) and B R,n (f n The numbers ) represent the interference airborne frequency f at the nth pulse. J,n The intermediate frequency bandwidth and radar carrier frequency are f n The intermediate frequency bandwidth.

[0027] The interference noise signal ratio (JNSR) of the jammer on the nth pulse is expressed as follows:

[0028]

[0029] Furthermore, step S2 is specifically as follows:

[0030] S21. Establish a generalized Markov decision process;

[0031] We assume the jammer is an intelligent agent and model the radar jamming process as a generalized Markov decision process. (Time-domain quintuple) Its specific definition in the time domain is as follows:

[0032] (1) Intelligent agent set The intelligent agent is composed of interference mechanisms.

[0033] (2) Action set The action set is formed by selecting the duration of interference generated by the defuzzification process after passing through the fuzzy inference system. The jammer's action a at time step n (i.e., the nth pulse) t,nRepresented by the pulse repetition period of the nth emitted pulse, i.e.

[0034] (3) State set Each rule in the fuzzy system rule base constitutes a state set.

[0035] (4) State transition probability P t :P t (s′|s,a) represents the jammer starting from state s t,n When =s, action a is executed. t,n =a, state transitions to s t,n+1 The transition probability of s′ is expressed as follows:

[0036] P t (s′|s,a)=P t (s t,n+1 =s′|s t,n =s,a t,n =a) (7)

[0037] Among them, P t (s′|s t ,a t It is considered unknown.

[0038] (5) Rewards Collection Treating the jamming duration and radar pulse repetition period generated by the jammer through fuzzy inference as random variables, the concept of relevant entropy is introduced to design the reward set. The specific design expression is as follows:

[0039]

[0040] Where σ represents the bandwidth of the positive definite kernel; T J T represents the coverage area of ​​the jammer's transmitted signal in the time domain; R This indicates the coverage area of ​​the radar signal in the time domain; in practice, This indicates that a specific amount of data is used to estimate the correlation entropy between two variables.

[0041] Five-tuple of a generalized Markov model in the frequency domain In the frequency domain, the specific definition is as follows:

[0042] (1) Intelligent agent set The intelligent agent is composed of interference mechanisms.

[0043] (2) Action set All selectable carrier frequencies of the jammers constitute the action set. The jammer's action a at time step n (i.e., the nth pulse) nRepresented by the carrier frequency of the nth transmitted pulse, i.e.

[0044] (3) State set All selectable carrier frequencies of the radar constitute the state set. The radar state s at time step n+1 n+1 Represented by the carrier frequency of the interference, i.e.

[0045] (4) State transition probability P f :P f (s′|s f ,a f ) indicates that the jammer has switched from state s f,n When =s, action a is executed. f,n =a, state transitions to s f,n+1 The transition probability of s′ is expressed as follows:

[0046] P f (s′|s,a)=P(s f,n+1 =s′|s f,n =s,a f,n =a) (9)

[0047] Among them, P f (s′|s f ,a f It is considered unknown.

[0048] (5) Rewards Collection The reward set consists of the JNSR of the jamming radar pulses. The reward for the jammer at time step n is represented as R. n =JNSR n JNSR n It is obtained from equation (6).

[0049] S22. Obtain the state-action value function;

[0050] Define the state value function generated by the agent in the time domain. The expression is as follows:

[0051]

[0052] in, This represents the input passed to the fuzzy inference system. The expression is as follows:

[0053]

[0054] in, This represents the defuzzification method for l rules in a fuzzy inference system. Represented as a fuzzy set F i l Membership degree, using the fuzzy Q algorithm to generate the time-domain global optimal action value function. The expression is as follows:

[0055]

[0056] Where M represents the number of rules in the fuzzy inference system, F i l This represents the calculated time error corresponding to the fuzzy set of the fuzzy variable. The expression is as follows:

[0057]

[0058] Where, r t+1 Let represent the reward value at time t+1, and γ∈[0,1] represent the discount rate. The update expression for the q function is as follows:

[0059]

[0060] in, This indicates that at time t+1, given rule l and agent behavior... The associated q function, This indicates that at time t, given rule l and agent behavior... The correlation q function, α t This represents the learning rate.

[0061] The optimal strategy obtained by setting the agent in the frequency domain The expression is as follows:

[0062]

[0063] in, The state-action value function, or Q-function, is defined as follows:

[0064]

[0065] Among them, E π [·] indicates the expected value, R n+k+1 γ represents the reward of the jammer at the (n+k+1)th time step, γ∈[0,1] represents the discount rate, k represents the kth time step after the nth time step, and γ k G represents the weighted value obtained from the discount rate. n This represents the discounted return at time step n, which is obtained by summing the future returns after weighting them by γ.

[0066] Furthermore, in step S3, the specific solution process for the joint interference algorithm of time-domain fuzzy Q-learning and frequency-domain Q-learning is as follows:

[0067] S31. Let i = 1, initialize the time-domain action set. Time-domain state set Time-domain discount rate γ t Temporal learning rate α t Initialize the fuzzy inference system, determine the number of input and output variables, the number of fuzzy labels, and the membership function; initialize the frequency domain action set. Frequency domain state set Frequency domain discount rate γ f Frequency domain learning rate α f ;

[0068] Wherein, the learning rate α t ,α f >0, its value decreases as the iteration proceeds.

[0069] S32. Set I coherent processing intervals CPI for training. Each CPI is considered as one iteration. At the beginning of the i-th (i = 1, 2, ..., I) iteration, randomly initialize the initial time-domain action a. t,0 Frequency domain action a f,0 Radar initial carrier frequency f r,0 Obtain the initial state s in the time domain t,0 Initial state s in the frequency domain f,0 ;

[0070] S33. For each pulse n (n = 0, 1, ..., L-1) in each CPI, select action a for each rule in the time domain according to the ε-greedy strategy. t,m , m = m + 1; in the frequency domain, select action a according to the ε-greedy strategy. f,n , n = n + 1;

[0071] S34. The radar signal pulse repetition period sensed from the environment is t. PRT,0 This is used as input to the fuzzy inference system to infer the time-domain global action selection 'a'. all,t,m The carrier frequency of the radar signal sensed from the environment is f. J,n ;

[0072] S35, Comprehensive Temporal Global Action Selection a all,t,m t PRT and r t Frequency domain a f,n f r,n and r f The time-frequency domain joint reward R is calculated. m,n R m,n =r t ·r f; respectively obtain new time-domain input values ​​t PRT,n and frequency domain state s f,n+1 =f r,n ;

[0073] In this context, (·) represents a multiplication operation.

[0074] S36. Update the time-domain q-function of the agents respectively. and frequency domain Q function

[0075]

[0076]

[0077] Where a, s, R, s′ represent s respectively n ,a f,n ,R m,n ,s f,n+1 .

[0078] S37, Update Status t,n =s t,n+1 s f,n =s f,n+1 If n < L, return to step S33; otherwise, proceed to step S38.

[0079] S38. Let i = i + 1. If i < I, return to step S34; otherwise, end the algorithm execution.

[0080] Each time the jammer interferes with a radar pulse, it outputs a time-frequency domain jamming strategy selection, ultimately achieving joint time-frequency domain jamming.

[0081] The beneficial effects of this invention are as follows: First, based on the radar signal processing flow, a target echo model and an interference signal model are established. The jammer is considered an intelligent agent, and the time-frequency domain joint interference process is modeled as a generalized Markov decision process, yielding a state value function. The proposed time-domain fuzzy Q-learning algorithm and frequency-domain Q-learning algorithm are used to jointly solve the problem, finally obtaining the jammer's interference time and carrier frequency selection strategy. This invention's method can select a suitable frequency band and duration for jamming the radar based on the radar's current frequency band and pulse repetition period, thereby avoiding interference with possible frequency bands at the next moment and determining the interference duration. By modeling the interference process as a generalized Markov decision process, the jammer's carrier frequency selection is considered an action in the frequency domain, and the radar carrier frequency is considered a state. The jammer's interference duration selection is considered an action in the frequency domain, and each fuzzy rule is considered a state. The product of the interference duration inferred from the time domain and the radar pulse repetition time's correlation entropy and the frequency domain interference JNSR is used as the joint reward function, achieving a time-frequency domain joint interference effect. Attached Figure Description

[0082] Figure 1 This is a flowchart of a time-frequency domain joint jamming method for frequency-agile radar based on fuzzy reinforcement learning, according to the present invention.

[0083] Figure 2 This is a flowchart of the time-frequency domain interference joint algorithm solution in an embodiment of the present invention.

[0084] Figure 3 The diagram shows the simulation results of time-domain JSR under various methods in the embodiments of the present invention.

[0085] Figure 4 The figure shows the results of simulating the number of time-domain interference pulses under various methods in the embodiments of the present invention.

[0086] Figure 5 This is a diagram showing the time-domain range of the radar signal during the final iteration of the simulation in this embodiment of the invention.

[0087] Figure 6 This is a diagram showing the time-domain range of the radar signal in the last iteration of the simulation using the random interference algorithm in this embodiment of the invention.

[0088] Figure 7 This is a diagram showing the time-domain coverage of the interference signal in the last iteration of the simulation using the time-frequency domain interference joint algorithm in this embodiment of the invention. Detailed Implementation

[0089] The method of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0090] like Figure 1 The flowchart of a time-frequency domain joint jamming method for frequency-agile radar based on fuzzy reinforcement learning according to the present invention is shown below. The specific steps are as follows:

[0091] S1. Based on the radar signal processing flow, establish the target echo model and the interference signal model;

[0092] S2. Treat the jammer as an intelligent agent and model the jamming process of the jammer jamming the frequency-agile radar as a generalized Markov decision process, and obtain the state value function of frequency domain Q-learning and the state value function of time domain fuzzy Q-learning respectively.

[0093] S3. Solve the generalized Markov decision problem in step S2 using the agent time-domain fuzzy Q-learning algorithm and the frequency-domain Q-learning algorithm to obtain the jammer's jamming duration selection strategy and jammer frequency domain selection strategy, ensuring that the jammer can achieve joint jamming in the time and frequency domains after the radar frequency changes.

[0094] In this embodiment, step S1 is specifically as follows:

[0095] The radar's transmitted signal is configured to undergo carrier frequency agility between pulses, with the number of pulses in a pulse train being L. N represents the set of selectable carrier frequencies for radar transmit pulses. R This indicates the total number of carrier frequencies that can be selected for radar transmission pulses, and sets that any two selectable carrier frequencies do not overlap.

[0096] in, The i-th element in is represented as f i =f i-1 +Δf R ,i∈{2,3,…,N R}, f i f represents the selection of the frequency value of the i-th element; i-1 Indicates the selection of the frequency value of the (i-1)th element; Δf R This represents a fixed frequency step value, with the radar bandwidth set to B. R The bandwidth for suppressing interference signals is B. J The relationship between the two is B. J =5~10B R The expression for the radar's transmitted pulse train signal x(t) is as follows:

[0097]

[0098] Where t represents time T r f represents the pulse repetition period. n Let represent the carrier frequency of the nth pulse of the radar, and s(t) represent the linear frequency modulated signal with unit amplitude.

[0099] When a radar transmits a signal to illuminate a target, generating a target echo, the echo signal y after the radar is affected by interference and noise at the nth pulse will be considered as follows. n The expression for (t) is as follows:

[0100]

[0101] Where, τ n =2R n / c indicates delay, R n Let f represent the radial distance between the radar and the target at the nth pulse, c represent the propagation speed of the electromagnetic wave, and f represent the radial distance between the radar and the target at the nth pulse. d,n N represents the Doppler frequency shift. n (t) represents a zero-mean power P. N Gaussian white noise, J n (t) represents the unit interference signal, β n α represents the change in amplitude caused by the propagation of the interference signal in space. n Let represent a parameter that incorporates propagation effects and target scattering; its amplitude expression is as follows:

[0102]

[0103] Among them, P R,n P represents the received power of the nth pulse at the radar receiver. T G represents the radar transmit power. T G R λ and σ represent the radar transmitting antenna gain and receiving antenna gain, respectively, λ represents the signal wavelength, and σ represents the target cross section (RCS).

[0104] A jammer is configured as an intelligent agent that uses fuzzy reinforcement learning in the frequency domain to generate a time-domain jamming duration strategy to interfere with radar signals. Upon receiving the radar signal, the jammer's receiver measures the pulse repetition period (PRT). Inevitably, there will be some error during the measurement process. If the time error between the jamming pulse and the frequency-agile radar pulse is large, it will severely affect the jamming effect; that is, when the radar frequency changes, the jammer's jamming time will be too long, causing it to lose synchronization with the radar. The measured PRT is input into the fuzzy inference system to generate a smaller PRT output value, which is the jammer's jamming duration, serving as the global action in the time domain.

[0105] The jammer is configured as an intelligent agent that generates a strategy in the frequency domain through reinforcement learning to jam radar signals. This indicates that the interference signal can be selected from a set of carrier frequencies.

[0106] Where, N J This indicates the number of selectable carrier frequencies for the interference signal. The p-th element in is represented as f J,p =f J,p-1 +Δf J , Δf J This represents a fixed interference frequency step value, where the interference carrier frequency is randomly or according to a certain strategy from... Select the appropriate option to set the frequency hopping range of the interference to cover all possible frequency bands of the radar system; the amplitude expression of the interference signal received by the radar is as follows:

[0107]

[0108] Among them, P I,n P represents the interference power at the radar receiver. J f represents the jamming transmission power. J,n Indicates the interfering carrier frequency, G J G represents the gain of the interfering transmitting antenna. R U represents the radar receiving antenna gain, and the probability of interference from the jammer to the radar. n The expression is as follows:

[0109]

[0110] Among them, B J,n (f J,n ) and B R,n (f n The numbers ) represent the interference airborne frequency f at the nth pulse. J,n The intermediate frequency bandwidth and radar carrier frequency are f n The intermediate frequency bandwidth.

[0111] The interference noise signal ratio (JNSR) of the jammer on the nth pulse is expressed as follows:

[0112]

[0113] In this embodiment, step S2 is specifically as follows:

[0114] S21. Establish a generalized Markov decision process;

[0115] We assume the jammer is an intelligent agent and model the radar jamming process as a generalized Markov decision process. (Time-domain quintuple) Its specific definition in the time domain is as follows:

[0116] (1) Intelligent agent set The intelligent agent is composed of interference mechanisms.

[0117] (2) Action set The action set is formed by selecting the duration of interference generated by the defuzzification process after passing through the fuzzy inference system. The jammer's action a at time step n (i.e., the nth pulse) t,n Represented by the pulse repetition period of the nth emitted pulse, i.e.

[0118] (3) State set Each rule in the fuzzy system rule base constitutes a state set.

[0119] (4) State transition probability P t :P t (s′|s,a) represents the jammer starting from state s t,n When =s, action a is executed. t,n =a, state transitions to s t,n+1 The transition probability of =s′ is expressed as:

[0120] P t (s′|s,a)=P t (s t,n+1 =s′|s t,n =s,a t,n=a) (7)

[0121] Among them, P t (s′|s t ,a t It is considered unknown.

[0122] (5) Rewards Collection Treating the jamming duration and radar pulse repetition period generated by the jammer through fuzzy inference as random variables, the concept of relevant entropy is introduced to design the reward set. The specific design expression is as follows:

[0123]

[0124] Where σ represents the bandwidth of the positive definite kernel; T J T represents the coverage area of ​​the jammer's transmitted signal in the time domain; R This indicates the coverage area of ​​the radar signal in the time domain; in practice, This indicates that a specific amount of data is used to estimate the correlation entropy between two variables.

[0125] Five-tuple of a generalized Markov model in the frequency domain In the frequency domain, the specific definition is as follows:

[0126] (1) Intelligent agent set The intelligent agent is composed of interference mechanisms.

[0127] (2) Action set All selectable carrier frequencies of the jammers constitute the action set. The jammer's action a at time step n (i.e., the nth pulse) n Represented by the carrier frequency of the nth transmitted pulse, i.e.

[0128] (3) State set All selectable carrier frequencies of the radar constitute the state set. The radar state s at time step n+1 n+1 Represented by the carrier frequency of the interference, i.e.

[0129] (4) State transition probability P f :P f (s′|s f ,a f ) indicates that the jammer has switched from state s f,n When =s, action a is executed. f,n =a, state transitions to s f,n+1 The transition probability of s′ is expressed as follows:

[0130] P f(s′|s,a)=P(s f,n+1 =s′|s f,n =s,a f,n =a) (9)

[0131] Among them, P f (s′|s f ,a f It is considered unknown.

[0132] (5) Rewards Collection The reward set consists of the JNSR of the jamming radar pulses. The reward for the jammer at time step n is represented as R. n =JNSR n JNSR n It is obtained from equation (6).

[0133] S22. Obtain the state-action value function;

[0134] Define the state value function generated by the agent in the time domain. The expression is as follows:

[0135]

[0136] in, This represents the input passed to the fuzzy inference system. The expression is as follows:

[0137]

[0138] in, This represents the defuzzification method for l rules in a fuzzy inference system. Represented as a fuzzy set F i l Membership degree, using the fuzzy Q algorithm to generate the time-domain global optimal action value function. The expression is as follows:

[0139]

[0140] Where M represents the number of rules in the fuzzy inference system, F i l This represents the calculated time error corresponding to the fuzzy set of the fuzzy variable. The expression is as follows:

[0141]

[0142] Where, r t+1 Let represent the reward value at time t+1, and γ∈[0,1] represent the discount rate. The update expression for the q function is as follows:

[0143]

[0144] in, This indicates that at time t+1, given rule l and agent behavior... The associated q function, This indicates that at time t, given rule l and agent behavior... The correlation q function, α t This represents the learning rate.

[0145] The optimal strategy obtained by setting the agent in the frequency domain The expression is as follows:

[0146]

[0147] in, The state-action value function, or Q-function, is defined as follows:

[0148]

[0149] Among them, E π [·] indicates the expected value, R n+k+1 γ represents the reward of the jammer at the (n+k+1)th time step, γ∈[0,1] represents the discount rate, k represents the kth time step after the nth time step, and γ k G represents the weighted value obtained from the discount rate. n This represents the discounted return at time step n, which is obtained by summing the future returns after weighting them by γ.

[0150] like Figure 2 As shown in this embodiment, the solution process of the time-domain fuzzy Q-learning and frequency-domain Q-learning joint interference algorithm (time-frequency joint interference algorithm) in step S3 is as follows:

[0151] S31. Let i = 1, initialize the time-domain action set. Time-domain state set Time-domain discount rate γ t Temporal learning rate α t Initialize the fuzzy inference system, determine the number of input and output variables, the number of fuzzy labels, and the membership function; initialize the frequency domain action set. Frequency domain state set Frequency domain discount rate γ f Frequency domain learning rate α f ;

[0152] Wherein, the learning rate α t ,α f >0, its value decreases as the iteration proceeds.

[0153] S32. Set I coherent processing intervals CPI for training. Each CPI is considered as one iteration. At the beginning of the i-th (i = 1, 2, ..., I) iteration, randomly initialize the initial time-domain action a. t,0 Frequency domain action a f,0 Radar initial carrier frequency f r,0 Obtain the initial state s in the time domain t,0 Initial state s in the frequency domain f,0 ;

[0154] S33. For each pulse n (n = 0, 1, ..., L-1) in each CPI, select action a for each rule in the time domain according to the ε-greedy strategy. t,m , m = m + 1; in the frequency domain, select action a according to the ε-greedy strategy. f,n , n = n + 1;

[0155] S34. The radar signal pulse repetition period sensed from the environment is t. PRT,m This is used as input to the fuzzy inference system to infer the time-domain global action selection 'a'. all,t,m The carrier frequency of the radar signal sensed from the environment is f. J,n ;

[0156] S35, Comprehensive Temporal Global Action Selection a all,t,m t PRT and r t Frequency domain a f,n f r,n and r f The time-frequency domain joint reward R is calculated. m,n R m,n =r t ·r f ; respectively obtain new time-domain input values ​​t PRT,n and frequency domain state s f,n+1 =f r,n ;

[0157] In this context, (·) represents a multiplication operation.

[0158] S36. Update the time-domain q-function of the agents respectively. and frequency domain Q function

[0159]

[0160]

[0161] Where a, s, R, s′ represent s respectively n ,a f,n ,R m,n ,s f,n+1 .

[0162] S37, Update Status t,n =s t,n+1 s f,n =s f,n+1 If n < L, return to step S33; otherwise, proceed to step S38.

[0163] S38. Let i = i + 1. If i < I, return to step S34; otherwise, end the algorithm execution.

[0164] When jamming a radar pulse, the jammer considers both the time and frequency domains. In the time domain, the radar signal collected by the receiver is measured using PRT. Each time a carrier frequency is selected, the globally optimal time-domain action is inferred based on the sensed jamming carrier frequency using fuzzy Q-learning, and Q-learning is used to find the Q-value. π (s,a) is used to obtain the optimal carrier frequency point. By combining the time-frequency domain reward value function, the joint time-frequency domain interference is finally achieved.

[0165] The present invention also provides another embodiment, which verifies and analyzes the method of the present invention through simulation:

[0166] During target detection by the frequency-agile radar, the radar transmits a linear frequency modulation (LFM) signal. After being reflected by the target, the signal is received by the jamming receiver, which then selects a time-frequency domain jamming strategy. The radar's frequency hopping range is set to [2,3] GHz, and its instantaneous bandwidth is 20 MHz. Its frequency sweeping strategy involves changing the frequency in steps with an 80% probability and randomly selecting the jamming frequency with a 20% probability. The frequency step value is Λf. J =20MHz. The number of pulses in one CPI transmitted by the radar is L=100, the transmit antenna gain and the receive antenna gain are both 30dB, and the pulse width is T. p =20μs, PRT is T r =100μs.

[0167] The initial target position is set directly above the radar at a distance of 100km, with a target RCS of 1m. 2 The target-borne jammer emits a frequency-hopping jamming signal with a frequency range of [2,3] GHz and an instantaneous bandwidth of B. J =200MHz, reinforcement learning parameters are set as follows: temporal learning rate is α t = 0.3 - 0.003n, where n represents the time step (i.e., the nth pulse), and the time-domain discount rate is γ. t =0.9, frequency domain learning rate α f =0.3-0.003n, where n represents the time step (i.e., the nth pulse), and the frequency domain discount rate is γ.t =0.9, and the total number of training iterations is 1000.

[0168] Figure 3 The figures show the simulation results of time-domain JSR under various methods in this embodiment of the invention. During the simulation, considering that radar typically uses multiple pulse coherent accumulation for signal processing, it is assumed that the radar performs signal processing every 10 consecutive pulses. That is, every 10 pulse repetition cycles, the radar signal appears randomly within the time observation window. The fuzzy Q-learning algorithm first uses Q-learning to track the time the radar signal appears within the observation window. After receiving the radar signal, it uses the fuzzy Q-learning algorithm to select the interference duration strategy, with the fuzzy inference system being the most important part. In contrast, the Q-learning algorithm only tracks the time the radar signal appears within the observation window and randomly selects the interference duration strategy after receiving the radar signal. It can be seen that the fuzzy Q-learning algorithm achieves a 5dB improvement in convergence compared to the Q-learning algorithm.

[0169] Figure 4 This is a graph showing the results of simulating the number of time-domain interference pulses under various methods in embodiments of the present invention. From Figure 4 It can be seen that random interference is ineffective due to the lack of tracking of the radar signal occurrence time and selection of interference duration. Within a single CPI (containing 100 radar pulses), only a dozen or so pulses are interfered with. Even after tracking the radar signal occurrence time using Q-learning, the lack of an interference duration strategy still results in low interference efficiency; after the algorithm stabilizes, only about forty pulses are interfered with within a single CPI (containing 100 radar pulses). In contrast, temporal fuzzy Q-learning not only tracks the radar signal occurrence time but also uses a fuzzy inference system to generate the interference duration strategy. Ultimately, after the algorithm stabilizes, seventy pulses are interfered with within a single CPI (containing 100 radar pulses). The proposed fuzzy Q-learning algorithm has the best interference performance among all compared methods. While maintaining the same interference capability, it determines the interference duration for the radar signal, thus achieving better interference results than the Q-learning interference algorithm that simply tracks the radar signal occurrence time.

[0170] Figure 5 This is a graph showing the time-domain range of the radar signal during the final iteration of the simulation in this embodiment of the invention. The graph illustrates the distribution of the radar signal in the time domain within the observation window.

[0171] Figure 6This is a graph showing the radar signal time-domain range result after the final iteration of the random interference algorithm in this embodiment of the invention. The darker areas represent the distribution of the radar signal in the time domain within the observation window, while the lighter areas represent the time range affected by the interference. It can be seen that the time-domain range of the radar signal covered by the interference is very small.

[0172] Figure 7 This is a graph showing the time-domain coverage of the interference signal during the final iteration of the time-frequency domain interference joint algorithm in this embodiment of the invention. The darker areas represent the distribution of radar signals within the observation window in the time domain, while the lighter areas represent the time range affected by the interference. It can be seen that the vast majority of radar signals are covered by the interference signal, representing a significant improvement compared to random interference.

[0173] In summary, the method of this invention can select a suitable frequency band and duration for jamming radar based on the radar's current frequency band and pulse repetition period. By modeling the jamming process as a generalized Markov decision process, the selection of the jammer's carrier frequency is considered an action in the frequency domain, and the radar carrier frequency is considered a state. In the time domain, the selection of the jammer's jamming duration is considered an action, and each fuzzy rule is considered a state. The product of the correlation entropy between the jamming duration inferred in the time domain and the radar pulse repetition time, and the frequency domain jamming JNSR is used as a joint reward function to achieve a joint time-frequency domain jamming effect.

[0174] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the implementation methods of the present invention, and should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of the present invention.

Claims

1. A time-frequency domain joint jamming method for frequency-agile radar based on fuzzy reinforcement learning, the specific steps of which are as follows: S1. Based on the radar signal processing flow, establish the target echo model and the interference signal model; S2. Treat the jammer as an intelligent agent and model the jamming process of the jammer jamming the frequency-agile radar as a generalized Markov decision process, and obtain the state value function of frequency domain Q-learning and the state value function of time domain fuzzy Q-learning respectively. S3. Solve the generalized Markov decision problem in step S2 using the agent time-domain fuzzy Q-learning algorithm and the frequency-domain Q-learning algorithm to obtain the jammer's jamming duration selection strategy and jammer frequency domain selection strategy, ensuring that the jammer can achieve joint jamming in the time and frequency domains after the radar frequency changes.

2. The time-frequency domain joint jamming method for frequency-agile radar based on fuzzy reinforcement learning according to claim 1, characterized in that, The specific steps of S1 are as follows: The radar's transmitted signal is configured to undergo carrier frequency agility between pulses, with one pulse train containing 100 pulses. ; This represents the set of carrier frequencies that can be selected for radar transmit pulses. This indicates the total number of selectable carrier frequencies for radar transmission pulses and sets that no two selectable carrier frequencies overlap. in, The first in Each element is represented as , Indicates the first Selecting the frequency value of each element; Indicates the first Selecting the frequency value of each element; This represents a fixed frequency step value, setting the radar bandwidth to [value]. The bandwidth for suppressing interference signals is The relationship between the two is: Radar transmit pulse train signal The expression is as follows: (1); in, Indicates time, Indicates the pulse repetition period. The radar's first The carrier frequency of each pulse A linear frequency modulated signal representing a unit amplitude; When the radar's transmitted signal illuminates the target and generates a target echo, the radar will then... The echo signal after being affected by interference and noise at each pulse point The expression is as follows: (2); in, Indicates time delay. Indicates the radar and target at the 1st Radial distance per pulse, Indicates the speed of propagation of electromagnetic waves. Indicates Doppler frequency shift, Let the power be zero mean. Gaussian white noise, Indicates unit interference signal, This indicates the change in amplitude caused by the propagation of the interference signal in space; Let represent a parameter that incorporates propagation effects and target scattering; its amplitude expression is as follows: (3); in, The radar's first The received power of each pulse at the radar receiver Indicates radar transmit power. These represent the radar transmitting antenna gain and receiving antenna gain, respectively. Indicates the signal wavelength. Represents the target's cross section (RCS); The jammer is set as an intelligent agent to generate a jamming duration strategy in the time domain through fuzzy reinforcement learning; the measured pulse repetition period (PRT) is input into the fuzzy inference system to generate a pulse repetition period output value with smaller error, which is the jammer's jamming duration, and is used as the global action selection in the time domain. The jammer is configured as an intelligent agent that generates a strategy in the frequency domain through reinforcement learning to jam radar signals. This indicates that the interference signal can be selected from a set of carrier frequencies; in, This indicates the number of selectable carrier frequencies for the interference signal. The first in Each element is represented as , This represents a fixed interference frequency step value, where the interference carrier frequency is randomly or according to a certain strategy from... Select the appropriate option to set the frequency hopping range of the interference to cover all possible frequency bands of the radar system; the amplitude expression of the interference signal received by the radar is as follows: (4); in, This indicates the interference power at the radar receiver. Indicates the jamming transmission power. Indicates the interfering carrier frequency. Indicates the gain of the interfering transmitting antenna. This represents the radar receiving antenna gain and the probability of interference from the jammer to the radar. The expression is as follows: (5); in, and They represent the first time. The interference airborne frequency is [frequency value] per pulse. The intermediate frequency bandwidth and radar carrier frequency are Intermediate frequency bandwidth; The jammer was in the first The expression for the interference noise signal ratio (JNSR) of each pulse is as follows: (6)。 3. The time-frequency domain joint jamming method for frequency-agile radar based on fuzzy reinforcement learning according to claim 1, characterized in that, Step S2 is as follows: S21. Establish a generalized Markov decision process; The jammer is treated as an intelligent agent, and the jamming radar process is modeled as a generalized Markov decision process; time-domain quintuple Its specific definition in the time domain is as follows: (1) Intelligent agent set The intelligent agent is composed of interference mechanisms. ; (2) Action set The action set is formed by selecting the duration of interference generated by the defuzzification process after passing through the fuzzy inference system. The jammer was in the first Time step action The launch of the The pulse repetition period of a pulse is represented as follows: ; (3) State set Each rule in the rule base of a fuzzy system constitutes a state set. ; (4) State transition probability : Indicates the jammer is in state Execute action at time The state transitions to The transition probability is expressed as follows: (7); in, Considered unknown; (5) Reward Set Treating the jamming duration and radar pulse repetition period generated by fuzzy inference as random variables, the concept of relevant entropy is introduced to design the reward set. The specific design expression is as follows: (8); in, Indicates the bandwidth of a positive definite core; This indicates the coverage area of ​​the jammer's transmitted signal in the time domain; This indicates the coverage area of ​​the radar signal in the time domain; in practice, This indicates that a specific amount of data is used to estimate the correlation entropy between two variables; Five-tuple of a generalized Markov model in the frequency domain In the frequency domain, the specific definition is as follows: (1) Intelligent agent set The intelligent agent is composed of interference mechanisms. ; (2) Action set The selectable carrier frequencies of all jammers constitute the action set. The jammer was in the first Time step action The launch of the The carrier frequency representation of each pulse, i.e. ; (3) State set All selectable carrier frequencies of the radar constitute the state set. ; Radar in the state of time step Represented by the carrier frequency of the interference, i.e. ; (4) State transition probability : Indicates the jammer is in state Execute action at time The state transitions to The transition probability is expressed as follows: (9); in, Considered unknown; (5) Reward Set The reward set consists of the JNSR of the jamming radar pulses. The jammer was in the first The reward at each time step is represented as , It is obtained from equation (6); S22. Obtain the state-action value function; Define the state value function generated by the agent in the time domain. The expression is as follows: (10); in, This represents the input passed to the fuzzy inference system. The expression is as follows: (11); in, This indicates that the fuzzy inference system targets the rule base. Methods for defuzzifying rules Represented as a fuzzy set Membership degree, using the fuzzy Q algorithm to generate the time-domain global optimal action value function. The expression is as follows: (12); in, This represents the number of rules in a fuzzy inference system. This represents the calculated time error corresponding to the fuzzy set of the fuzzy variable. The expression is as follows: (13); in, express The reward value at any moment, Indicates the discount rate; The function's update expression is as follows: (14); in, Indicates in Time-based given rules and agent behavior The following associations function, Indicates in Time-based given rules and agent behavior The following associations function, Indicates the learning rate; The optimal strategy obtained by setting the agent in the frequency domain The expression is as follows: (15); in, The state-action value function, or Q-function, is defined as follows: (16); in, This indicates the calculation of mathematical expectation. Indicates the jammer at the 1st Rewards at each time step Indicates the discount rate. Indicates the first After the first time step Each time step This represents the weighted value obtained from the discount rate. Indicates the first The discounted return at the time step is determined by future returns. We obtain the weighted sum.

4. The time-frequency domain joint jamming method for frequency-agile radar based on fuzzy reinforcement learning according to claim 3, characterized in that, In step S3, the specific solution process for the joint interference algorithm of time-domain fuzzy Q-learning and frequency-domain Q-learning is as follows: S31, Order Initialize the time-domain action set Time-domain state set Time-domain discount rate Time-domain learning rate ; Initialize the fuzzy inference system, determine the number of input and output variables, the number of fuzzy labels, and the membership function; initialize the frequency domain action set. Frequency domain state set Frequency domain discount rate Frequency domain learning rate ; Among them, learning rate Its value decreases as the iteration proceeds; S32, Set up co-training There are n coherent processing intervals CPI, each CPI is considered as one iteration, and in the nth iteration... At the start of the next iteration, the initial time-domain actions are randomly initialized. Frequency domain action Radar initial carrier frequency Obtain the initial state in the time domain Initial state of the frequency domain ; S33, For each pulse in each CPI In the time domain, according to A greedy strategy selects an action for each rule. , In the frequency domain, according to Greedy strategy selects action , ; S34. The radar signal pulse repetition period sensed from the environment is... This is used as input to the fuzzy inference system to infer the global action selection in the time domain. The carrier frequency of radar signals sensed from the environment is ; S35, Comprehensive Temporal Global Action Selection , and Frequency domain , and The joint reward in the time and frequency domains was calculated. , ; respectively obtain new time-domain input values and frequency domain state ; in, Indicates multiplication operation; S36. Update the time domain of the agents respectively. function and frequency domain function ; (17); (18); in, They represent ; S37, Update Status , ,when If the condition is met, return to step S33; otherwise, proceed to step S38. S38, Order ,like Return to step S34; otherwise, terminate the algorithm execution. Each time the jammer interferes with a radar pulse, it outputs a time-frequency domain jamming strategy selection, ultimately achieving joint time-frequency domain jamming.

Citation Information

Patent Citations

  • Deep reinforcement learning air combat game interpretation method and system based on fuzzy decision tree

    CN111353606A

  • Radar anti-interference intelligent decision-making method based on backtracking Q learning

    CN115508790A