A radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method

By employing a multi-domain joint intelligent active anti-jamming method that integrates radar space, time, frequency, and energy, and utilizing the Q-learning algorithm to optimize radar resource scheduling, the anti-jamming challenges of modern radar under various electronic interference conditions are solved, thereby improving the radar's survivability and mission performance.

CN115932750BActive Publication Date: 2026-03-27UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-23
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing radar anti-jamming methods are insufficient to effectively cope with modern flexible electronic jamming, especially in the absence of prior information. Single-domain parameter scheduling cannot fully leverage the advantages of multi-functional radar resources, resulting in insufficient anti-jamming capabilities.

Method used

A multi-domain joint intelligent active anti-jamming method based on radar space-time-frequency-energy is adopted. By establishing an adversarial scenario between the surveillance radar and the jammer, the Q-learning algorithm is used to optimize the multi-domain resource scheduling strategy of the radar, such as wave position, frequency, power and dwell time, to dynamically avoid interference.

Benefits of technology

It significantly enhances the radar's survivability in electronic warfare environments, improves mission performance, taps into the radar's potential for anti-jamming, and enables coordinated control of multi-domain resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115932750B_ABST
    Figure CN115932750B_ABST
Patent Text Reader

Abstract

The application discloses a radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method, first establishes a confrontation scene under the radar and multi-jammer environment, then sorts out the model parameters and working modes of the radar and the jammer, determines the evaluation index to complete the evaluation of the radar scheduling strategy, and then converts the problem into a Markov decision process, and uses Q learning to solve the optimal strategy of radar multi-domain resource scheduling. The optimization strategy obtained by the method has strong adaptability and good model scalability, and can be adjusted according to the actual need of the scheduled parameter resource. Through the collaborative control of the transmission parameters between the multiple domains of the radar, compared with the existing single-domain resource scheduling anti-jamming method, the radar active anti-jamming capability can be significantly enhanced while improving the task performance, and the potential task performance of the radar is excavated. Meanwhile, Q learning is used to optimize the radar node multi-domain joint active anti-jamming strategy, and the survival ability of the radar in the electronic countermeasure environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of electronic countermeasures, and particularly relates to a radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method. BACKGROUND

[0002] Modern new electronic jamming mechanisms (active repeater jamming, deception jamming, smart jamming) are rich in types, increase in jamming means, and enhance in jamming strength, which puts forward new requirements for radar detection and survival, and the radar needs to further develop corresponding anti-jamming means to cope with active electronic jamming in the environment.

[0003] In the past 20 years, radar anti-jamming means have been widely researched, and from the perspective of signal processing, such as adaptive filtering and blind source separation algorithm, are also applied to jamming suppression. This kind of method is that after the radar is subjected to electronic jamming, passive measures are taken according to the type of signal jamming, which is called passive anti-jamming. On the one hand, the jamming strength is enhanced, and once the radar is jammed, the passive anti-jamming means is difficult to completely offset the influence of the jamming on the signal; on the other hand, due to the multiple types and flexible and changeable ways of electronic jamming, a single form of passive anti-jamming means is difficult to cope with complex and combined jamming.

[0004] In view of the limitations of passive anti-jamming in the actual electronic countermeasure environment, active anti-jamming is proposed. The connotation of active anti-jamming is to avoid the attack of electronic jamming on the radar by adjusting the radar parameters before the radar is subjected to electronic jamming. Unlike passive anti-jamming, because it is carried out before being jammed, active anti-jamming lacks prior information about the environment and jamming, and the direction of scheduling radar resources to implement active anti-jamming is not clear enough.

[0005] And reinforcement learning can overcome the pain point of active anti-jamming, because it does not rely on prior information, has strong environmental adaptability, and can effectively guide the radar active anti-jamming under the condition of lack of prior information. The literature "A Reinforcement Learning Based Approach for Multitarget Detection in Massive MIMO Radar, IEEE Transactions on Aerospace and Electronic Systems, vol. 57, no. 5, pp. 2622-2636" proposes to establish a beam forming model of large-scale MIMO radar based on reinforcement learning. Through optimizing the spatial resource, the algorithm can detect targets with lower signal-to-noise ratio. But the model considered is relatively single, only considering the optimization scheduling of MIMO radar spatial resource; In the literature "Radar active antagonism through deep reinforcement learning: A Way to address the challenge of mainlobe jamming, Signal Processing, 2021, 186: 108130" proposes to establish a model of radar frequency hopping to avoid main lobe interference, but this model only focuses on the strategy optimization in the frequency domain. The existing active anti-jamming methods only aim at the dynamic configuration of single domain parameter / resource, without considering further expanding the dimension of resource scheduling. In view of the multi-domain parameter agility of modern multi-function radars such as MIMO and digital array radars, only scheduling single domain transmission parameters is difficult to take advantage of the diversity of modern radar resources. SUMMARY

[0006] To solve the above technical problems and further improve the survivability in the radar electronic countermeasure environment, the present application provides a radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method.

[0007] In order to facilitate the description of the content of the present application, the following terms are first explained:

[0008] Term 1: ELINT system

[0009] The ELINT system refers to an electronic intelligence reconnaissance system (ELINT), which can receive and analyze signals of radar and other radiation sources.

[0010] Term 2: wave position

[0011] Wave position is a discrete expression of radar beam pointing parameters. For example: assuming that the radar azimuth angle is 0-20°, the azimuth angle beam width is 5°, and the entire azimuth space is to be covered, at least four wave positions are needed.

[0012] The technical scheme of the present application is: a radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method, the specific steps are as follows:

[0013] Step S1, establish the confrontation scene under the monitoring radar and the multi-jammer environment, determine the number of radars and jammers and the confrontation relationship;

[0014] Step S2, sort out the specific resource parameters that the monitoring radar can jointly schedule;

[0015] Step S3, sort out the time-varying parameters of the jammer, and establish the jammer interception mechanism;

[0016] Step S4, initialize the radar system and jammer parameters, and calculate the detection performance and low interception performance;

[0017] Step S5, according to the detection probability, low interception performance and interception penalty, build rewards, and convert the problem into a Markov decision process;

[0018] Step S6, use the Q-learning algorithm to solve the Markov decision process, and obtain the optimized scheduling strategy.

[0019] Further, the step S1 is specifically as follows:

[0020] Based on the frequency agility radar of pulse compression, all wave positions in the scanned target space are periodically scanned, there are n jammers in the space, it is determined that the radar and the jammer are in confrontation relationship, and under the premise of meeting the detection performance, the radar dynamically schedules its own resources to avoid the interference of the jammer.

[0021] Further, the step S2 is specifically as follows:

[0022] The time for the radar to scan each wave position is recorded as a scheduling interval SI, in each scheduling interval, the transmission parameters of the radar remain unchanged, the radar uses the selected parameters to detect externally, and the detection performance and low interception performance are calculated according to the detection result.

[0023] Assuming that the target space is divided into p wave positions, and the divided wave positions are numbered, the radar needs to consume p scheduling intervals to complete a complete periodic scan. In each scheduling interval, the radar selects a wave position from the wave position set Θ:

[0024] Θ={Θ1,Θ2,Θ3...Θ p} (1)

[0025] The radar selects a frequency from a frequency set F, which has q available frequencies in total:

[0026] F = {f1, f2, f3... f q} (2)

[0027] The total frequency band is equally divided, and the relationship between the tth frequency and the t-1th frequency can be expressed by equation (3):

[0028] f t = f t-1 + δf t∈2,3,..q (3)

[0029] Where δf represents a fixed frequency interval step.

[0030] The radar power selectable set P is shown in equation (4), and the power has u selectable values. The dwell time selectable set T d is shown in equation (5), and the dwell time has v selectable values:

[0031] P = {P1, P2, P3... P u} (4)

[0032] T d = {T d1 , T d2 , T d3 ... T dv} (5)

[0033] Further, the step S3 is specifically as follows:

[0034] The jammer knows all the frequencies that the radar may use through the ELINT system in advance. The jammer randomly scans all q frequency points in the frequency set, and the instantaneous coverage frequency band is updated in a fixed time slot T mf , and then a new round of periodic scanning is started when the jammer scans all q frequency points.

[0035] The spatial position of the jammer is time-varying, Θ j represents the wave position of the jammer, R j represents the distance of the jammer from the radar, and the position of the jammer at a given time k can be represented as When the jammer intercepts the radar signal, it performs noise suppression jamming, otherwise, the jammer remains silent to save energy loss.

[0036] Further, in the step S4, the basic parameters of the radar system and the jammer are initialized, and the radar performance index is calculated, which is specifically as follows:

[0037] S41, calculate the detection probability P d ;

[0038] If the radar signal is not jammed, according to the radar equation, the signal-to-noise ratio ξ o is obtained by the radar receiver

[0039]

[0040] where P i represents the power selected by the radar, G r represents the main lobe gain of the radar beam, λ represents the wavelength of the radar signal, σ represents the target scattering cross section area, R represents the distance between the radar and the target, L represents the radar loss in dB, K = 1.38 * 10 -23 , represents the Boltzmann constant, T e represents the effective noise temperature, F r represents the internal noise factor of the receiver, B r represents the radar operating bandwidth. If the radar is jammed by a jammer, the jamming power P jr received by the radar is

[0041]

[0042] where G(θ) is a function of angle θ, representing the receiving beam gain of the radar when the jamming energy interferes from different azimuth angles of the radar beam; it is assumed that when the jamming enters from the main lobe of the radar beam, G(θ) = G r , and when it enters from the side lobe, G(θ) = G sav , G sav represents the average side lobe gain of the radar antenna, P j represents the transmitted signal power of the jammer, G j represents the beam gain of the jammer, then the signal-to-jamming-and-noise ratio ζ o obtained by the radar receiver is

[0043]

[0044] Without jamming, first pulse compression is performed on the echo, and the signal-to-noise ratio ξ o(pc) after pulse compression is calculated

[0045] ξ o(pc) = ξ o D = ξ o τB r (9)

[0046] where D represents the time-bandwidth product of the radar, τ represents the pulse width of the radar, and the pulse compression process improves the signal-to-noise ratio of the original pulse by D times. The radar stays at each wave position for a dwell time T di , which determines the number of pulses that can be accumulated at the current wave position, and the number of pulses is

[0047]

[0048] wherein the symbol denotes the floor function, T r denotes the pulse repetition interval. Rewrite ξ o(pc) as ξ1, calculate the accumulated signal-to-noise ratio ξ p after non-coherent accumulation of n nci pulses:

[0049]

[0050] Based on the calculation of formula (9) to formula (11), the detection probability is obtained:

[0051]

[0052] wherein Q denotes the Marcum Q function, P fa denotes the false alarm rate.

[0053] An approximate calculation can be made for formula (12):

[0054]

[0055] In the presence of interference, the subsequent calculation process is similar to the case without interference, and ξ o is replaced by ζ o , which can be calculated by formula (8) to formula (11).

[0056] S42, calculate the low intercept performance P ni ;

[0057] The low intercept performance of radar for each scheduling is calculated, and the single pulse intercept probability P I can be represented as:

[0058] P I = P I_f · P I_e · P I_t · P I-s · P I_ρ (14)

[0059] wherein P I denotes the probability of a single pulse being intercepted, P I_f denotes the probability of the frequency domain being intercepted, P I_e denotes the probability of the energy domain being intercepted, P I_t denotes the probability of the time domain being intercepted, P I_s denotes the probability of the space domain being intercepted, and P I_ρ denotes the probability of the polarization domain being intercepted.

[0060] The interceptor operates in frequency sweep mode. The probability of being intercepted in the frequency domain is equivalent to the probability of the radar signal frequency band entering the instantaneous bandwidth of the interceptor. Therefore, P I f for:

[0061]

[0062] Among them, B ins B represents the instantaneous intercept bandwidth of the radar interceptor. total This indicates the total frequency band range of the radar.

[0063] Assuming the interceptor is always in receiving mode, the probability of being intercepted in the time domain is numerically equal to the duration of the radar signal emitted per unit time, i.e., the duty cycle:

[0064]

[0065] Energy domain interception probability P I_e That is, the detection probability of the interceptor. If the radar signal is higher than the detection threshold at the output of the interceptor, it will be detected with probability P. I_e Successfully intercepted:

[0066]

[0067] Where, γ j B represents the polarization mismatch factor. j L represents the bandwidth of the jamming signal from the jammer. j Indicates the jammer's interception loss, F j This indicates the noise figure inside the receiver intercepted by the jammer.

[0068] Assuming the interceptor receiver has sidelobe detection capability, meaning it can receive radar pulses omnidirectionally, then P I_s =1, and temporarily disregarding the polarization domain factor, let P be denoted as P. I_ρ =1, then equation (14) simplifies to:

[0069]

[0070] Define the minimum number of pulses to be intercepted, n. least =35, indicating that the interceptor must continuously intercept at least 35 radar pulses to be considered a successful interception. Therefore, the probability of intercepting fewer than 35 pulses can be defined as the probability P of not being intercepted under a multi-pulse system. ni , using P ni To measure the low intercept performance of a radar:

[0071]

[0072] in, This represents the number of combinations.

[0073] Further, the step S5 is specifically as follows:

[0074] The radar resource scheduling problem is converted into a Markov decision process, and the radar is in a state s i selects an action a i , and completes state transition to s i+1 , and the interference environment feeds back an immediate reward r(s i , a i+1 ) according to a i+1 and s i . First, the radar state is defined, and the radar needs to cover all wave positions in the airspace within a scanning period. It is assumed that the scanning of each wave position needs a scheduling interval SI, and the state s can be defined as the sequence number of the scheduling interval:

[0075] s i ∈{S∣S={1,2,3...p}} (20)

[0076] Wherein, S represents the set of states, s i represents the i-th state, and the state s is increased by 1 every scheduling interval, and when a new round of scanning period starts, the state s is reset to 1.

[0077] In each scheduling interval, the radar selects an action a i from the action set A to complete dynamic management of the radar resource, and the action of the radar is composed of wave position, frequency, power and residence time, and the wave position cannot be repeatedly selected within the same period:

[0078] a i ∈{A∣A=(Θ,F,P,T d )} (21)

[0079] When the radar completes the detection of a wave position according to the parameter combination of the selected action, the interference environment will feed back an immediate reward r according to the detection result. The reward can quantify the effectiveness of this radar scheduling, and the reward is composed of three parts, i.e., detection reward low-interception reward and interception penalty r p . If the radar signal is intercepted by multiple jammers at the same time, the interception penalties of multiple jammers need to be superimposed, as shown in equation (26):

[0080]

[0081]

[0082]

[0083]

[0084]

[0085] Each action will get the immediate reward r corresponding to the action, and the total cumulative reward of a scan cycle is obtained by accumulating the immediate reward of the cycle:

[0086]

[0087] Further, the step S6 is specifically as follows:

[0088] The Q-learning uses an epsilon-greedy strategy π to select actions:

[0089] a i ~ π (· | s i )(28)

[0090] Wherein, ε represents a decaying random probability, and ε ∈ [0, 1], and the epsilon-greedy strategy selects the action currently considered to be the maximum behavior value with a probability of 1-ε, and selects an action a from all m selectable actions with a probability of ε:

[0091]

[0092] Wherein, Q represents the state behavior value in reinforcement learning, and the Q table is updated by using a greedy strategy μ:

[0093] a i+1 ~ μ (· | s i+1 )(30)

[0094] Then, according to the tuple data (s i , a i , s i+1 , a i+1 , r) obtained by one action interaction, the Q value is updated according to the following formula:

[0095]

[0096] Wherein, α represents a learning rate, and γ represents a decay factor.

[0097] With the updating of the Q table, after sufficient exploration in the action space, the Q table and the corresponding epsilon-greedy strategy converge to obtain the optimal action strategy of this decision-making process, that is, the optimal solution of the scheduling problem.

[0098] The method of the application first establishes the confrontation scene under the radar and multi-jammer environment, then sorts out the model parameters and working mode of the radar and jammer, determines the evaluation index to complete the evaluation of the radar scheduling strategy, and then converts the problem into a Markov decision process, and uses Q learning to solve the optimal strategy of radar multi-domain resource scheduling. The optimization strategy obtained by the method of the application has strong adaptability and good model scalability, and can be adjusted according to the actual need of the scheduled parameter resources. By cooperatively regulating the transmission parameters between the multiple domains of the radar, compared with the existing single-domain resource scheduling anti-jamming method, the radar active anti-jamming capability can be significantly enhanced while improving the task performance, and the potential task performance of the radar is excavated. At the same time, Q learning is used to optimize the radar node multi-domain joint active anti-jamming strategy, and the survival ability of the radar in the electronic countermeasure environment is improved. BRIEF DESCRIPTION OF DRAWINGS

[0099] Figure 1 A flow chart of a radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method of the application.

[0100] Figure 2 A schematic diagram of a monitoring radar and multi-jammer confrontation scene in an embodiment of the application.

[0101] Figure 3 A schematic diagram of a monitoring radar multi-domain resource scheduling in an embodiment of the application.

[0102] Figure 4 A schematic diagram of a jammer parameter time-varying process and confrontation result analysis in an embodiment of the application.

[0103] Figure 5 A schematic diagram of radar and jammer environment interaction in an embodiment of the application.

[0104] Figure 6 A convergence curve diagram of Q learning radar cumulative return in an embodiment of the application.

[0105] Figure 7 A comparison diagram of the cumulative return of Q learning radar and random radar in the first to the 100th round in an embodiment of the application.

[0106] Figure 8 A comparison diagram of the cumulative interception times of Q learning radar and random radar in the first to the 100th round in an embodiment of the application. DETAILED DESCRIPTION

[0107] The application mainly adopts a simulation experiment method for verification, and all steps and conclusions are verified correct on python3.7. The application will be further described below in combination with the drawings and embodiments.

[0108] As Figure 1As shown, a radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method flow chart of the present application, the specific steps are as follows:

[0109] Step S1, establish the confrontation scene under the monitoring radar and multi-jammer environment, determine the number of radars and jammers and the confrontation relationship;

[0110] Step S2, sort out the specific resource parameters that can be jointly scheduled by the monitoring radar;

[0111] Step S3, sort out the time-varying parameters of the jammer, and establish the jammer interception mechanism;

[0112] Step S4, initialize the radar system and jammer parameters, and calculate the detection performance and low interception performance;

[0113] Step S5, according to the detection probability, low interception performance and interception penalty, build the reward, and convert the problem into a Markov decision process;

[0114] Step S6, use the Q-learning algorithm to solve the Markov decision process, and get the optimized scheduling strategy.

[0115] In this embodiment, the step S1 is specifically as follows:

[0116] Based on the frequency agility radar of pulse compression, all wave positions in the scanned target space domain are periodically scanned, there are n jammers in the space domain to interfere, it is determined that the radar and the jammer are in confrontation relationship, and under the premise of meeting the detection performance, the radar dynamically schedules its own resources to avoid the interference of the jammer.

[0117] In this embodiment, the step S2 is specifically as follows:

[0118] The time for the radar to scan each wave position is recorded as the scheduling interval SI, in each scheduling interval, the transmission parameters of the radar remain unchanged, the radar uses the selected parameters to detect externally, and the detection performance and low interception performance are calculated according to the detection result.

[0119] Assuming that the target space domain is divided into p wave positions, and the divided wave positions are numbered, the radar needs to consume p scheduling intervals to complete a complete periodic scan. In each scheduling interval, the radar selects a wave position from the wave position set Θ:

[0120] Θ={Θ1,Θ2,Θ3...Θ p} (1)

[0121] The wave position cannot be selected repeatedly in the same period, and the optimization of wave position selection can significantly improve the success rate of avoiding interference.

[0122] The radar selects a frequency from the frequency set F, and there are q available frequencies in the set:

[0123] F = {f1, f2, f3...f q} (2)

[0124] The total frequency band is equally divided, and the relationship between the tth frequency and the t-1th frequency can be expressed by equation (3):

[0125] f t = f t-1 + δf t∈2,3,..q (3)

[0126] Where δf represents a fixed frequency interval step.

[0127] The radar power and the dwell time affect the calculation of the received echo signal-to-noise ratio, and the dwell time determines the number of echo pulses that can be accumulated in this wave position. The radar power set P is selected as shown in equation (4), and the power has u selectable values. The dwell time set T d is selected as shown in equation (5), and the dwell time has v selectable values:

[0128] P = {P1, P2, P3...P u} (4)

[0129] T d = {T d1 ,T d2 ,T d3 ...T dv} (5)

[0130] In this embodiment, the interference environment parameters, the Q algorithm parameters, and the radar parameters are shown in Table 1, Table 2, and Table 3, respectively.

[0131] Table 1

[0132]

[0133] Table 2

[0134] Parameter Symbol Value Learning rate α 0.01 Decay factor Gamma 0.9 Decayed random probability Epsilon 0.8~0.02 Training times e num ]]> 500

[0135] Table 3

[0136]

[0137] In this embodiment, the step S3 is specifically as follows:

[0138] The jammer learns all the frequencies that the radar can use through the ELINT system in advance, and the jammer randomly scans all q frequency points in the frequency set, and the instantaneous coverage frequency band is updated in a fixed time slot T mf When the jammer scans all q frequency points, it starts a new round of periodic scanning.

[0139] The spatial position of the jammer is time-varying, Θ j represents the wave position where the jammer is located, R j represents the distance from the radar to the jammer, given a time k, the position of the jammer can be represented as When the jammer intercepts the radar signal, it performs noise suppression jamming, otherwise, the jammer remains silent to save energy loss.

[0140] In this embodiment, in the step S4, the basic parameters of the radar system and the jammer are initialized, and the radar performance index is calculated, as follows:

[0141] S41, calculate the detection probability P d ;

[0142] If the radar signal is not jammed, according to the radar equation, the signal-to-noise ratio ξ o received by the radar receiver is:

[0143]

[0144] where P i represents the power selected by the radar, G r represents the main lobe gain of the radar beam, λ represents the wavelength of the radar signal, σ represents the target scattering cross-sectional area, R represents the distance between the radar and the target, L represents the radar loss, K = 1.38 * 10 -23 , represents the Boltzmann constant, T e represents the effective noise temperature, F r represents the receiver internal noise coefficient, B r represents the radar operating bandwidth. If the radar is jammed by the jammer, the interference power P jr received by the radar receiver is:

[0145]

[0146] where G(θ) is a function of angle θ, representing the gain of the receiving beam of the radar when the interference energy interferes from different azimuth angles of the radar beam; it is assumed that when the interference enters from the main lobe of the radar beam, G(θ) = G r , when it enters from the side lobe, G(θ) = G sav , G sav represents the average side lobe gain of the radar antenna, P j represents the signal power transmitted by the jammer, G j represents the beam gain of the jammer, then the signal-to-interference-and-noise ratio ζ o received by the radar receiver is:

[0147]

[0148] Firstly, the echo is pulse compressed without interference, and the signal-to-noise ratio ξ after pulse compression is calculated o(pc) :

[0149] ξ o(pc) =ξ o D=ξ o τB r (9)

[0150] Wherein, D represents the time bandwidth product of the radar, τ represents the pulse width of the radar, and the pulse compression process increases the signal-to-noise ratio of the original pulse by D times. The radar stays for a dwell time T at each wave position di , the dwell time determines the number of pulses that can be accumulated at the current wave position, and the number of pulses is:

[0151]

[0152] Wherein, the symbol represents the floor function, T r represents the pulse repetition interval. Rewrite ξ o(pc) as ξ1, calculate the accumulated signal-to-noise ratio ξ p after incoherent accumulation of n nci pulses:

[0153]

[0154] Based on the calculation of formula (9) to formula (11), the detection probability is calculated:

[0155]

[0156] Wherein, Q represents the Marcum Q function, P fa represents the false alarm rate.

[0157] The formula (12) can be calculated approximately:

[0158]

[0159] Under the condition of interference, the subsequent calculation process is similar to that under the condition of no interference, and ξ o under the condition of no interference is replaced with ζ o , which can be calculated by formula (8) to formula (11).

[0160] S42, calculate the low intercept performance P ni ;

[0161] The low intercept performance of the radar is calculated every time the radar is scheduled. Assuming that the single pulse intercept probability P I can be represented as:

[0162] P I =P I-f ·PI-e • P I-t • P I-s • P I_ρ (14)

[0163] where P I represents the probability of single pulse interception, P I_f represents the probability of frequency domain interception, P I_e represents the probability of energy domain interception, P I_t represents the probability of time domain interception, P I_s represents the probability of space domain interception, P I_ρ represents the probability of polarization domain interception.

[0164] The interception receiver works in the sweep frequency mode, and the probability of frequency domain interception is equivalent to the probability that the radar signal frequency band falls into the instantaneous bandwidth of the interceptor, so P I_f is:

[0165]

[0166] where B ins represents the instantaneous interception bandwidth of the radar interceptor, B total represents the total frequency range of the radar.

[0167] The time domain interception probability refers to the probability that the interception receiver is in the receiving state when the signal energy just reaches the front end of its antenna. Assuming that the interception receiver is always in the receiving state, the time domain interception probability is numerically equal to the time of the radar radiation signal in a unit of time, i.e. the duty cycle:

[0168]

[0169] The energy domain interception probability P I_e is the detection probability of the interception receiver. If the radar signal is higher than the detection threshold at the output end of the interception receiver, it will be successfully intercepted with a probability of P I_e :

[0170]

[0171] where γ j represents the polarization mismatch factor, B j represents the jamming signal bandwidth of the jammer, L j represents the interception loss of the jammer, F j represents the internal noise coefficient of the jammer interception receiver.

[0172] Assuming that the interception receiver has the ability of sidelobe detection, i.e. it can receive radar pulses omnidirectionally, P I_s = 1, and the polarization domain factor is not considered for the time being, denoted as P I_ρ = 1, then formula (14) is simplified as:

[0173]

[0174] After the interceptor acquires a signal, it needs to measure the signal parameters, sort and identify the signal. Acquiring only a single radar pulse is insufficient to complete subsequent signal processing tasks. Furthermore, defining an acquisition as a single radar pulse results in low reliability. The minimum number of intercepted pulses, n, is defined as follows. least =35, indicating that the interceptor must continuously intercept at least 35 radar pulses to be considered a successful interception. Therefore, the probability of intercepting fewer than 35 pulses can be defined as the probability P of not being intercepted under a multi-pulse system. ni , using P ni To measure the low intercept performance of a radar:

[0175]

[0176] in, This represents the number of combinations.

[0177] In this embodiment, step S5 is specifically as follows:

[0178] The radar resource scheduling problem is transformed into a Markov decision process, where the radar is in state s. i Next select action a i Complete the state transition to s i+1 Interference environment according to a i and s i+1 Feedback with an immediate reward r(s) i+1 ,a i State, action, and reward are the core elements of a Markov decision process. First, define the radar state. The radar needs to cover all airspace positions within one scan cycle. Assuming that scanning each position requires one scheduling interval SI, the state s can be defined as the sequence number of the scheduling interval:

[0179] s i ∈{S∣S={1,2,3...p}} (20)

[0180] Where S represents the set of states, s i This represents the i-th state. The state s is incremented by 1 after each scheduling interval. When a new round of scanning begins, the state s is reset to 1.

[0181] At each scheduling interval, the radar selects an action a from action set A. i Dynamic management of radar resources is achieved, with radar operations consisting of wave position, frequency, power, and dwell time. It is important to note that the wave position cannot be selected repeatedly within the same period.

[0182] a i∈ {A | A = (Θ, F, P, T d )} (21)

[0183] When the radar completes the detection of a wave position according to the parameter combination of the selected action, an instant reward r is fed back by the jamming environment according to the detection result. The reward can quantify the effectiveness of this radar scheduling, and the reward is composed of three parts, i.e., a detection reward a low-interception reward and an interception penalty r p . The detection reward and the low-interception reward are simulated by using a step function, and a detection performance threshold and a low-interception performance threshold are defined. When the set threshold is exceeded, a positive reward is given to the sub-item, and otherwise, a non-positive reward is given. The interception penalty is used to distinguish the punishment degree of the main lobe interception and the sidelobe interception in the numerical value, and the punishment of the main lobe interception needs to be greater than that of the sidelobe interception. If the radar signal is intercepted by multiple jammers at the same time, the interception penalties of multiple jammers need to be superimposed, as shown in equation (26):

[0184]

[0185]

[0186]

[0187]

[0188]

[0189] An instant reward r corresponding to the action is obtained for each action, and the total cumulative reward in a scanning period is obtained by accumulating the instant rewards in the scanning period:

[0190]

[0191] In the embodiment, the step S6 is specifically as follows:

[0192] An action is selected by using an e-greedy strategy π of Q-learning:

[0193] a i ~ π (· | s i ) (28)

[0194] Wherein, ε represents a probability value set in advance, and ε ∈ [0, 1], and the e-greedy strategy selects an action a that is currently considered to have the maximum behavior value with a probability of 1-ε, and selects an action a from all m selectable actions at random with a probability of ε:

[0195]

[0196] Wherein, Q represents the state-action value in reinforcement learning, and the Q table is updated by using the greedy strategy μ (always selecting the action with the current maximum action value):

[0197] a i+1 ~ μ (· | s i+1 )(30)

[0198] Then, the Q value is updated according to the tuple data (s i ,a i ,s i+1 ,a i+1 ,r) obtained by one action interaction according to the following formula:

[0199]

[0200] Wherein, α represents the learning rate, and γ represents the decay factor.

[0201] With the updating of the Q table, after sufficient exploration in the action space, the Q table and the corresponding ε-greedy strategy tend to converge, and the optimal action strategy of the decision-making process is obtained, that is, the optimal solution of the scheduling problem is obtained.

[0202] As Figure 2 shown, the monitoring radar and multi-jammer countermeasure scene schematic diagram in the embodiment of the application is that a single radar countermeasures three jammers, the radar position is fixed, and the jammers move in the airspace. Figure 3 The monitoring radar multi-domain resource scheduling schematic diagram is shown. At each scheduling time, the radar selects the wave position, frequency, power, and residence time parameters.

[0203] Figure 4 The jammer parameter time-varying process and countermeasure result analysis schematic diagram is shown. The receiver of the jammer works in the frequency sweeping mode, and the instantaneous interception frequency band and the wave position of the jammer change with time. The countermeasures of the radar and the jammer mainly have three possibilities, which are (1) not interfered, the radar wave position and the frequency band do not coincide with any jammer; (2) sidelobe interference, the radar wave position does not coincide, but the frequency band falls into the interception frequency band of a jammer, and when the energy detection requirement is met, the radar is intercepted by the sidelobe; (3) main lobe interception, the radar wave position and the frequency band coincide with a jammer, and the radar is intercepted by the main lobe.

[0204] Figure 5 The radar and jammer environment interaction schematic diagram is shown, Figure 6 The convergence curve graph of the cumulative return of the Q-learning radar with the increase of the number of training rounds is shown, Figure 7 The comparison graph of the cumulative returns of the Q-learning radar and the random radar in the first to the 100th round is shown, Figure 8 The comparison graph of the cumulative interception times of the Q-learning radar and the random radar in the first to the 100th round is shown, and the main lobe interception and the sidelobe interception are distinguished.

[0205] In summary, the method of the present application can well realize the strategy optimization of radar multi-domain resource joint active anti-jamming, enhance the active anti-jamming performance of the radar, and improve the survivability of the radar in the electronic countermeasure environment.

[0206] Those skilled in the art will understand that the embodiments described herein are for the purpose of helping the reader understand the principles of the present application and should be understood as not limiting the scope of protection of the present application to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations according to the technical inspiration disclosed in the present application without departing from the essence of the present application, and these modifications and combinations are still within the scope of protection of the present application.

Claims

1. A radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method, the specific steps being as follows: Step S1, establishing a confrontation scene under a monitoring radar and a multi-jammer environment, determining the number of radars and jammers and the confrontation relationship; Step S2, sorting out the specific resource parameters of the monitoring radar that can be jointly scheduled; Step S3, sorting out the time-varying parameters of the jammer, and establishing a jammer interception mechanism; Step S4, initializing the radar system and jammer parameters, and calculating the detection performance and low-interception performance; Step S5, constructing a reward according to the detection probability, low-interception performance and interception penalty, and converting the problem into a Markov decision process; Step S6, using a Q-learning algorithm to solve the Markov decision process to obtain an optimized scheduling strategy.

2. The radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method according to claim 1, characterized in that, The step S1 is specifically as follows: Based on the frequency agility radar of pulse compression, all wave positions in the scanned target space domain are periodically scanned, there are n jammers in the space domain to interfere, it is determined that the radar and the jammer are in a confrontation relationship, and under the premise of meeting the detection performance, the radar dynamically schedules its own resources to avoid the interference of the jammer.

3. The radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method according to claim 2, characterized in that, The step S2 is specifically as follows: The time for the radar to scan each wave position is recorded as a scheduling interval SI, in each scheduling interval, the transmission parameters of the radar remain unchanged, the radar uses the selected parameters to detect externally, and the detection performance and low-interception performance are calculated according to the detection results; It is assumed that the target space domain is divided into p wave positions, and the divided wave positions are numbered, then the radar needs to consume p scheduling intervals to complete a complete periodic scan; in each scheduling interval, the radar selects a wave position from the wave position set Θ: (1); The radar selects a frequency from the frequency set F, there are q available frequencies in the set: (2); The total frequency band is equally divided, then the relationship between the tth frequency and the t-1th frequency can be expressed by formula (3): (3); Wherein, δf represents a fixed frequency interval step; The radar power selectable set P is shown in equation (4), and the power has u selectable values. The dwell time selectable set T is also shown in equation (4). d See equation (5), the dwell time has v possible values: (4); (5)。 4. The radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method according to claim 3, characterized in that, The step S3 is specifically as follows: The jammer knows all the frequencies that the radar can use in advance through the ELINT system, and the jammer randomly scans all q frequency points in the frequency set, and the instantaneous coverage frequency range is in a fixed time slot T mf Update, when the jammer scans all q frequency points, a new round of periodic scanning is started again; The spatial position of the jammer is time-varying, Θ j represents the wave position where the jammer is located, R j represents the distance from the radar to the jammer, given a time k, the position of the jammer can be represented as (Θ j k , R j k ), when the jammer intercepts the radar signal, it carries out noise suppression jamming, otherwise, the jammer remains silent to save energy loss.

5. The radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method according to claim 4, characterized in that, In the step S4, the basic parameters of the radar system and the jammer are initialized, and the radar performance indicators are calculated, which are specifically as follows: S41, calculate the detection probability P d ; If the radar signal is not jammed, the signal-to-noise ratio (SNR) obtained by the radar receiver is, according to the radar equation, given by: S = 4PTG2λ2R4Tf (6); wherein, G represents the power selected by the radar r G represents the main lobe gain of the radar beam, λ represents the wavelength of the radar signal, σ represents the target scattering cross section, R represents the distance between the radar and the target, L represents the radar loss in dB, K = 1.38 * 10 -23 , represents the Boltzmann constant, T e F represents the effective noise temperature, F r B represents the receiver internal noise factor, B r B represents the radar operating bandwidth; if the radar is interfered by the jammer, the interference power received by the radar is: (7); where is a function of the angle θ and represents the gain of the radar's receive beam when the jammer energy is coming from different azimuth angles of the radar beam; it is assumed that when the jammer is coming from the main lobe of the radar beam, = G r when coming from the side lobes, = G sav , G sav represents the average side lobe gain of the radar antenna, P j represents the jammer transmitted signal power, G j represents the jammer beam gain, then the signal to jammer and noise ratio at the radar receiver is : (8); First, the echo is pulse compressed without interference, and the signal-to-noise ratio after pulse compression is calculated : (9); Wherein, represents the time-bandwidth product of the radar, τ represents the pulse width of the radar, and the pulse compression process increases the signal-to-noise ratio of the original pulse by D times; the radar stays for a stay time T at each wave position di , the stay time determines the number of pulses that can be accumulated at the current wave position, and the number of pulses is: (10); where the symbol represents the floor function, T r represents the pulse repetition interval; re-write is , the non-coherent accumulation n p is calculated after the accumulation signal-to-noise ratio : (11); On the basis of formula (9) to formula (11), the detection probability is solved: (12); where Q denotes the Marcum Q function, P fa represents the false alarm rate; The formula (12) can be approximately calculated: (13); With interference, the subsequent calculation process is similar to that without interference. (The sentence about interference-free calculations is incomplete and lacks context.) use The substitution can be performed using equations (8) to (11); S42, calculate low probability of intercept P ni ; The low probability of intercept performance of each radar schedule is calculated, assuming a single pulse intercept probability P I may be represented as: (14); where P I represents the probability of interception of the single pulse, P I_f represents the probability of interception of the frequency domain, P I_e represents the probability of interception of the energy domain, P I_t represents the probability of interception of the time domain, P I_s represents the probability of interception of the space domain, P I_ρ represents the probability of interception of the polarization domain; The intercept receiver works in the sweep mode, the probability of frequency domain being intercepted is equivalent to the probability of the radar signal frequency band falling into the instantaneous bandwidth of the intercept receiver, so P I_f is: (15); where B ins represents the instantaneous intercept bandwidth of the radar intercept receiver, B total represents the total frequency band range of the radar; It is assumed that the interception receiver is always in a receiving state, then the time domain interception probability is equal in value to the time of the radar radiation signal in a unit time, that is, the duty cycle: (16); The probability of intercept P I_e The probability of intercept P is the probability that a radar signal will be successfully intercepted by an intercept receiver if the radar signal is above the detection threshold of the output of the intercept receiver. I_e The probability of intercept P is the probability that a radar signal will be successfully intercepted by an intercept receiver if the radar signal is above the detection threshold of the output of the intercept receiver. (17); where γ j represents the polarization mismatch factor, B j represents the jammer jamming signal bandwidth, L j represents the jammer intercept loss, F j represents the jammer intercept receiver internal noise figure; Assuming the intercepted receiver has the ability of sidelobe detection, i.e. omni-directional receiving radar pulses, then P I_s =1, and temporarily not considering the polarization domain factor, let P I_ρ =1, then formula (14) is simplified as: (18); Definition of minimum number of intercepted pulses n least = 35, which means that the receiver must have intercepted at least 35 radar pulses consecutively to be considered a successful interception; then, on this basis, the probability of intercepting less than 35 pulses can be defined as the probability of not being intercepted in a multi-pulse system P ni , which measures the low-interception performance of the radar: ni ​ (19); Wherein, ℂ represents the combination number.

6. The radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method according to claim 5, characterized in that, The step S5 is specifically as follows: The radar resource scheduling problem is transformed into a Markov decision process, where the radar is in a state s i selects an action a i , and the state transitions to s i+1 , the interference environment feeds back an immediate reward r(s i+1 , a i ) according to a i and s i+1 . First, the radar state is defined. The radar needs to cover all the wave positions in the airspace within a scan period. It is assumed that the scan of each wave position takes one scheduling interval SI. The state s can be defined as the sequence number of the scheduling interval: (20); wherein, a set of states, represents the i-th state, and each time a scheduling interval passes, the state s is incremented by 1, and when a new scanning cycle is started, the state s is reset to 1. At each scheduling interval, the radar selects an action a from the action set A i The dynamic management of radar resources is completed, the action of the radar is composed of wave position, frequency, power and residence time, and the wave position cannot be repeatedly selected in the same period. (21); When the radar completes the detection of a wave position according to the parameter combination of the selected action, an instant reward r is fed back according to the detection result; the reward can quantify the effectiveness of this radar scheduling, and the reward is composed of three parts, i.e., a detection reward , a low-interception reward , and an interception penalty ; if the radar signal is intercepted by multiple jammers at the same time, the interception penalties of the multiple jammers need to be superimposed, as shown in formula (26): (22); (23); (24); (25); (26); Each action will obtain the immediate reward r corresponding to the action, and the total cumulative reward under the scanning period is obtained by accumulating the immediate rewards of each action: (27)。 7. The radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method according to claim 6, characterized in that, The step S6 is specifically as follows: The Q-learning uses an ε-greedy strategy π to select an action: (28); where ε denotes a decaying random probability, and The ε-greedy policy greedily chooses the action currently believed to have the maximum behavioral value with probability 1 - ε, and randomly chooses an action from all m available actions with probability ε : (29); Wherein, Q represents the state-action value in reinforcement learning, and the Q table is updated by using a greedy strategy μ: (30); Then the Q value is updated according to the tuple data (s i , a i , s i+1 , a i+1 , r) as follows: (31); wherein denotes a learning rate, denotes a decay factor; With the updating of the Q table, after sufficient exploration in the action space, the Q table and the corresponding ε-greedy strategy tend to converge, the optimal action strategy of the decision process is obtained, and the optimal solution of the scheduling problem is obtained.

Citation Information

Patent Citations

  • Deep reinforcement learning anti-interference method for frequency agile radar

    CN114509732A

  • Radar detection and tracking

    WO2022130350A1