Networking radar jamming and mutual jamming elimination method based on decentralized q learning
By constructing a multi-radar cooperative anti-interference and anti-mutual interference process based on decentralized Q-learning, and optimizing the frequency agility strategy, the problem of interference and mutual interference in multi-radar cooperative detection is solved, thereby improving the radar's detection performance and signal processing gain.
Patent Information
- Application Number
- CN202410140146.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-01-30
AI Technical Summary
In complex interference environments, mutual interference exists in multi-radar cooperative detection. Existing technologies are unable to effectively suppress interference and mutual disturbance, which affects detection performance.
By employing a decentralized Q-learning-based approach, the multi-radar collaborative anti-jamming and anti-interference process is constructed as a generalized Markov decision process. The frequency domain resource scheduling of the radar is optimized through a frequency agility strategy, and the DQJIE algorithm is used to solve the frequency agility strategy, thereby realizing intelligent frequency band selection and real-time interactive feedback among radars.
It effectively suppressed internal interference within the system, improved the detection performance of the networked radar, enhanced signal processing gain, and optimized the radar's anti-interference and anti-interference capabilities.
Smart Images

Figure CN118151104B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of radar anti-jamming, and particularly relates to a networking radar jamming and mutual jamming elimination method based on decentralized Q learning. BACKGROUND
[0002] Multi-radar cooperative detection technology has attracted extensive attention in recent years due to its good detection performance and anti-jamming performance. Compared with a single radar, multi-radar can cooperatively transmit electromagnetic signals, construct a high-dimensional signal observation space through integrated resource scheduling, and thus can cope with complex electromagnetic interference environments. With the increase in the number of radars, mutual interference may occur among the radars due to the scarcity of frequency spectrum resources, which will seriously restrict the radar detection performance. Therefore, it is an urgent need in the future to study the mutual interference suppression technology of multi-radar systems in complex interference environments.
[0003] In a complex interference environment, multiple radars are needed to cooperatively detect and track enemy targets. To better cope with complex environments and ensure detection performance, a dynamic and intelligent interference and mutual interference suppression method based on reinforcement learning is proposed. The process of multi-radar cooperative anti-jamming and anti-interference is constructed as a generalized Markov decision process, and the decentralized Q-learning method is used to obtain the multi-radar frequency agility strategy. A radar frequency agility method based on reinforcement learning is proposed in the literature“K. Li, B. Jiu, H. Liu, and S. Liang, Reinforcement learning based anti-jamming frequency hopping strategies design for cognitive radar, 2018 IEEE International Conference on Signal Processing, Communications and Computing (ICSPCC), Qingdao, China, 018, pp. 1-5.” to ensure the detection performance of the radar in an unknown interference environment. The radar learns the frequency variation strategy of external interference and internal mutual interference through interaction with the environment, thereby realizing the selection of the optimal frequency hopping strategy of the radar. A multi-agent cooperative anti-jamming method based on Markov game is proposed in the literature“F. Yao and L. Jia, A collaborative multi-agent reinforcement learning anti-jamming algorithm in wireless networks, IEEE Wireless Communications Letters, vol. 8, no. 4, pp. 1024-1027, Aug. 2019.” to solve the influence of external frequency sweeping interference and internal mutual interference on the transmission quality of the communication channel in the wireless communication network. The user interacts with the interference while exchanging information between users, realizes the selection of the optimal channel, and finally suppresses the interference and mutual interference. In these two methods, one is a single-station radar against interference and mutual interference in unknown interference information, and the other is a user in a communication network that coordinates to resist interference and mutual interference. At present, there is little research on the mutual interference problem in multi-radar cooperative detection in a complex interference environment. Most of the research is to eliminate interference or mutual interference independently, and focuses on solving such problems in communication systems or vehicle-mounted radar systems. Therefore, it is necessary to study the joint interference and mutual interference suppression method in the multi-radar cooperative scenario. SUMMARY
[0004] To solve the above technical problems, the application provides a networking radar interference and mutual interference elimination method based on decentralized Q-learning.
[0005] The technical scheme adopted by the application is as follows:
[0006] S1, according to the phased array radar signal processing flow, a target echo signal, interference signal and mutual interference signal model is established;
[0007] S2, the conditions of radar interference and mutual interference are analyzed, and a multi-radar frequency domain resource scheduling optimization model based on the SIJNR criterion is established according to the definition of the frequency aiming probability;
[0008] S3, each radar is regarded as an intelligent agent, and the multi-radar cooperative anti-interference and anti-interference process is constructed as a generalized Markov decision process, and the multi-radar frequency domain resource scheduling optimization model based on the SIJNR criterion in step S2 is converted into a value function optimization model;
[0009] S4, the interference and mutual interference suppression algorithm based on decentralized Q learning is used to solve the dynamic multi-agent frequency domain resource scheduling problem in step S2, and the frequency agility strategy of each radar is obtained.
[0010] Further, the step S1 is specifically as follows:
[0011] N radars working in self-generation and self-reception mode are set to cooperatively detect a target, and the radars work in time synchronization mode for communication.
[0012] Each radar transmits a linear frequency modulation signal, and the carrier frequency of the signal transmitted by each radar is agile between pulses, so that the transmission signal form of the i-th (i=1, 2,..., N) radar is represented as:
[0013]
[0014] Where t represents the time, M i represents the number of pulses transmitted by radar i in a coherent processing interval, T r represents the pulse repetition time, f i,m represents the carrier frequency of the m+1th transmission pulse of radar i, which can be represented as f i,m =f i,c +b m Δf, f i,c represents the initial carrier frequency of radar i, Δf represents the frequency hopping interval, b m represents the frequency coding sequence; s(t) represents the baseband linear frequency modulation signal, s m (t) = s(t-mT r ).
[0015] The radar i transmits signals to probe the target. Assuming that the interference signals and the mutual interference signals entering through the side lobe can be suppressed by the side lobe cancellation technology, the radar i will receive the scattered echo signals from the target, the interference signals and the mutual interference signals in the main lobe. The echo signals received by the radar i are represented as:
[0016]
[0017] wherein the five terms on the right side of the equation (2) represent the target scattered echo signals of the radar i, the mutual interference signals scattered by the radar k through the target, the direct wave mutual interference signals of the radar l, the interference signals and the noise signals respectively; τ i = 2(R i -v i t) / c represents the two-way time delay caused by the distance between the radar i and the target, R i represents the distance between the radar i and the target, v i represents the radial velocity of the target relative to the radar i, and c represents the speed of light; τ I,k,i = (R k + R i -v k,i t) / c represents the time delay caused by the signals transmitted by the radar k and scattered by the target to reach the radar i, v k,i = v k +v i represents the velocity of the target relative to the radar k and the radar i in the path of the radar i, τ D,l,i = R l,i / c represents the one-way time delay caused by the distance R l,i between the radar l and the radar i; f d,i = 2v i / λ i represents the Doppler frequency shift of the target relative to the radar i, f d,k,i = 2v k,i / λ i represents the Doppler frequency shift of the target relative to the radar k and the radar i in the path of the radar i; λ i represents the wavelength of the signals of the radar i; α i , β i , γ I,k,i and γ D,l,i represent the complex parameters related to the propagation loss; the amplitudes of the target echo signals, the interference signals, the scattered echo mutual interference signals and the direct wave mutual interference signals are represented as:
[0018]
[0019]
[0020]
[0021]
[0022] where P T,i , G T,i , G R,i and L s represent the peak transmit power, transmit antenna gain, receive antenna gain and system loss of radar i respectively; λ k represents the radar k signal wavelength, λ l represents the radar l signal wavelength; σ i represents the radar cross section of the target relative to radar i; P J , G T,J and λ J represent the peak transmit power, transmit antenna gain and wavelength of the jammer respectively.
[0023] The echo signal received by radar i is down-converted to the expression as follows:
[0024]
[0025] where, represents the jamming signal after mixing, represents the noise signal after mixing, represent the frequency components of the target scattering echo mutual interference signal and direct wave mutual interference signal respectively.
[0026] The mixed signal is low-pass filtered with radar i receiver bandwidth B i , and only when the jamming signal frequency band, mutual interference signal frequency band and radar i transmit signal frequency band have overlap, the jamming signal and mutual interference signal will enter the correlation processor of radar i.
[0027] Further, the step S2 is specifically as follows:
[0028] S21, analyze the conditions of radar frequency domain interference and mutual interference;
[0029] The frequency aiming probability of the jamming to the m+1th pulse of radar i is defined as:
[0030]
[0031] where ∩ represents the intersection operator, ΔB i and ΔB J represent the intermediate frequency bandwidth of radar i receiver and the instantaneous bandwidth of the jammer respectively, f J,m represents the center frequency of the jammer. Similarly, the frequency aiming probability of the radar k transmit signal to the m+1th pulse of radar i mutual interference is defined as:
[0032]
[0033]
[0034] where ΔB i and ΔB k denote the intermediate frequency bandwidth of radar i receiver and the instantaneous bandwidth of radar k transmit signal, respectively. and denote the frequency pointing probability of the scatter echo mutual interference signal and the direct wave mutual interference signal, respectively. denote the pointing probability of radar k transmit signal to radar i main lobe, which is expressed as follows:
[0035]
[0036] The interference signal and the mutual interference signal will cause interference to radar i if and only if the frequency pointing probability to radar i is not 0.
[0037] S22, establish a multi-radar frequency domain resource scheduling optimization model based on the SIJNR criterion;
[0038] The SIJNR definition expression of the m+1th pulse of radar is as follows:
[0039]
[0040] where P S,i,m , P J,i,m and P N,i denote the actual power of the target echo signal received by radar i at the m+1th pulse, the actual power of the interference signal and the noise signal power, respectively. I,k,i,m and P D,k,i,m denote the actual power of the scatter echo mutual interference signal and the direct wave mutual interference signal received by radar i at the m+1th pulse, respectively; and and γ D,k,i denote the complex amplitude parameter of the radar k transmit direct wave mutual interference signal received by radar i.
[0041] According to the radar SIJNR model, the frequency domain resource scheduling model is constructed as:
[0042]
[0043] where, denote the set of radar i selectable frequency points.
[0044] Further, the step S3 is specifically as follows:
[0045] S31, establish a generalized Markov decision process;
[0046] Each radar is regarded as an agent, and the process of multi-radar cooperative anti-jamming and anti-interference is constructed as a generalized Markov decision process, which is represented by a five-tuple :
[0047] (1) Agent set N transversal frequency radars constitute the agent set.
[0048] (2) Action set represents the joint action space of N agents, that is, the optional frequency range of all radars.
[0049] wherein a i,m = f i,m represents the action of radar i at the mth time step, a m = [f 1,m , f 2,m ,..., f N,m ] T represents the joint action of agents at the mth time step, (·) T represents the transpose operator.
[0050] (3) State set represents the entire optional jamming frequency range of the jammer, s m = f J,m represents the state presented by the jammer at the mth time step.
[0051] (4) State transition probability P: p(s′|s,a) represents the probability of the agent transitioning to state s′ under the action a in state s, and its definition is as follows:
[0052] p(s′|s,a) = P{s m = s′|s m-1 = s, a m-1 = a} (14)
[0053] (5) Reward set represents the joint reward space of N agents, represents the reward set of agent i under the strategy π, and the SIJNR of each radar at the mth time step constitutes the reward set, that is, r i,m (a i,m , a -i,m ) = SIJNR i,m ;
[0054] wherein a -i,m represents the joint action of other agents except agent i.
[0055] S32, a value function optimization model is established;
[0056] Rewrite as:
[0057]
[0058] wherein, denotes an expectation operator, γ ∈ [0, 1] denotes a discount factor, π i : s→a i denotes the policy of the agent i.
[0059] Rewrite as: a value function optimization model can be obtained:
[0060]
[0061] wherein, Q i (s, a i , a -i ) denotes the action value function of the agent i, π i (s, a i ) denotes the probability distribution of the action a i of the agent i under the policy π i .
[0062] Further, in the step S4, the DQJIE algorithm solving process is specifically as follows:
[0063] S41, set the common training I phase-locked intervals CPI, let j = 1, regard each CPI as a scene, initialize the discount factor γ ∈ [0, 1], the learning rate η ∈ (0, 1), the state space the action space the Q table Q i (s, a i , a -i ) of each agent and the global Q table Q(s, a), start the iteration of each scene;
[0064] S42, let m = 0, regard each pulse in each CPI as each time step, initialize s0 = f J,0 ;
[0065] S43: let i = 1, start the iteration of each step for the agent i;
[0066] S44, select the action a i,m for the m+1th pulse of the agent i based on the ε-greedy policy;
[0067] S45, based on the action a i,m of the agent i and the current time interference carrier frequency s m = fJ,m Obtain reward r i,m , and obtain the interference carrier frequency s of next time m+1 = f J,m+1 ;
[0068] S46, update the Q function Q of the agent i i (s, a i , a -i ):
[0069]
[0070] Wherein, s = s m , a i = a i,m , a -i = a -i,m , r i = r i,m , Indicates the set of nodes having a communication link with the agent i, Indicates the number of elements in the set ;
[0071] S47, the agent i shares its Q table Q i (s, a i , a -i ) to its neighbor nodes, that is, the nodes having a communication link with the agent i;
[0072] S48, let i = i + 1, if i ≤ N, return to step S44, otherwise execute step S49;
[0073] S49, update the global Q function Q (s, a):
[0074]
[0075] Wherein, s = s m , a i = a i,m , a -i = a -i,m .
[0076] S410, let m = m + 1, if m < M, return to step S43, otherwise execute step S411;
[0077] S411, let j = j + 1, if j ≤ I, return to step S42, otherwise execute step S412;
[0078] S412, output the Q table Q of each agent i (s, a i , a -i ) and the global Q table Q (s, a).
[0079] The method of the present application firstly establishes target echo signal, interference signal and mutual interference signal model according to phased array radar signal processing flow; secondly analyzes the conditions of radar interference and mutual interference, and establishes a multi-radar frequency domain resource scheduling optimization model based on signal-to-interference-and-noise ratio criterion; then the radar is regarded as an intelligent agent, the multi-radar cooperative anti-interference and anti-mutual interference process is constructed as a generalized Markov decision process, and the frequency domain resource scheduling optimization model is converted into a value function optimization model; finally, the interference and mutual interference suppression algorithm based on decentralized Q learning is used to solve the problem, and the frequency agility strategy of each radar is obtained. The method of the present application constructs the multi-radar cooperative anti-interference and anti-mutual interference process as a generalized Markov decision process, regards the interference frequency band as the state, the radar frequency band as the action, and the SIJNR as the reward function, each radar intelligently selects the radar frequency band according to the current interference frequency band, and the action is interacted and fed back in real time through the communication link based on the graph theory model between radars, and the Q table of each radar is corrected, so as to effectively suppress the mutual interference in the system while resisting the frequency sweeping interference, and effectively improve the detection performance of the networked radar. BRIEF DESCRIPTION OF DRAWINGS
[0080] Figure 1 A flow chart of the networked radar interference and mutual interference elimination method based on decentralized Q learning of the present application.
[0081] Figure 2 A networked radar cooperative detection schematic diagram in the embodiment of the present application.
[0082] Figure 3 A DQJIE algorithm flow schematic diagram in the embodiment of the present application.
[0083] Figure 4 A DQJIE algorithm solving flow chart in the embodiment of the present application.
[0084] Figure 5 A radar interference pulse number and mutual interference pulse number result chart under the simulation comparison method in the embodiment of the present application.
[0085] Figure 6 A radar mutual interference pulse number result chart under the simulation comparison method in the embodiment of the present application.
[0086] Figure 7 A radar pulse non-coherent processing SIJNR result chart under the simulation comparison method in the embodiment of the present application. DETAILED DESCRIPTION
[0087] The method of the present application will be further described below in combination with the drawings and embodiments.
[0088] As Figure 1As shown, a flow chart of a networking radar jamming and mutual interference elimination method based on decentralized Q-learning of the present application is shown, and the specific steps are as follows:
[0089] S1, according to the phased array radar signal processing flow, establish target echo signal, jamming signal and mutual interference signal model;
[0090] S2, analyze the conditions of radar jamming and mutual interference, and establish a multi-radar frequency domain resource scheduling optimization model based on the SIJNR criterion according to the definition of the frequency aiming probability;
[0091] S3, regarding each radar as an intelligent agent, constructing the multi-radar cooperative anti-jamming and anti-mutual interference process as a generalized Markov decision process, and converting the multi-radar frequency domain resource scheduling optimization model based on the SIJNR criterion in step S2 into a value function optimization model;
[0092] S4, using the decentralized Q-learning based jamming and mutual interference elimination (DQJIE) algorithm to solve the dynamic multi-agent frequency domain resource scheduling problem in step S2, and obtaining the frequency agility strategy of each radar.
[0093] In this embodiment, the step S1 is specifically as follows:
[0094] Suppose that N radars working in self-generation and self-reception work cooperatively to detect targets, and the radar nodes work in time synchronization mode for communication, and the scene diagram is as shown in Figure 2 .
[0095] Suppose that each radar transmits a linear frequency modulation (LFM) signal, and the carrier frequency of the signal transmitted by each radar is agile between pulses, and the form of the transmission signal of the i-th (i=1, 2,..., N) radar can be expressed as:
[0096]
[0097] Where t represents the time, M i represents the number of pulses transmitted by radar i within a coherent processing interval (CPI), T r represents the pulse repetition time (PRT), f i,m represents the carrier frequency of the m+1th transmission pulse of radar i, which can be expressed as f i,m =f i,c +b mΔf, f i,c represents the initial carrier frequency of radar i, Δf represents the frequency hopping interval, b m represents the frequency coding sequence; s(t) represents a baseband linear frequency modulation signal, which can be expressed in the form of:
[0098]
[0099] wherein T p represents the signal pulse width, μ=B / T p represents the frequency modulation slope, B represents the signal bandwidth, rect(·) represents the rectangular function, s m (t) = s(t-mT r ).
[0100] The radar i transmits a signal to detect a target. Assuming that the interference signal and the mutual interference signal entering through the side lobe can be suppressed through the side lobe cancellation technology, the radar i will receive the scattered echo signal from the target, the interference signal and the mutual interference signal through the main lobe. The echo signal received by the radar i can be expressed as:
[0101]
[0102] wherein the five terms on the right side of equation (3) respectively represent the target scattered echo signal of the radar i, the mutual interference signal after the radar k is scattered by the target, the direct wave mutual interference signal of the radar l, the interference signal and the noise signal; τ i = 2(R i -v i t) / c represents the two-way time delay caused by the distance between the radar i and the target, R i represents the distance between the radar i and the target, v i represents the radial velocity of the target relative to the radar i, c represents the speed of light; τ I,k,i = (R k +R i -v k,i t) / c represents the time delay generated when the signal transmitted by the radar k is scattered by the target and reaches the radar i, v k,i = v k +v i represents the velocity of the target relative to the radar k and the radar i, τ D,l,i = R l,i / c represents the one-way time delay caused by the distance R l,i between the radar l and the radar i; f d,i = 2v i / λ i represents the Doppler shift of the target relative to the radar i, f d,k,i = 2v k,i / λ idenotes the Doppler shift of the target relative to radar k receiving radar i transmitting path; λ i denotes the radar i signal wavelength; α i , β i , γ I,k,i and γ D,l,i denote the complex parameters related to propagation loss; the amplitudes of the target echo signal, jamming signal, scattering echo-jamming signal and direct wave-jamming signal can be respectively expressed as:
[0103]
[0104]
[0105]
[0106]
[0107] wherein P T,i , G T,i , G R,i and L s denote the peak transmit power, transmit antenna gain, receive antenna gain and system loss of radar i respectively; λ k denotes the radar k signal wavelength, λ l denotes the radar l signal wavelength; σ i denotes the radar cross section (RCS) of the target relative to radar i; P J , G T,J and λ J denote the peak transmit power, transmit antenna gain and wavelength of the jammer respectively.
[0108] The echo signal received by radar i is down-converted to obtain the expression as follows:
[0109]
[0110] wherein, denotes the jamming signal after mixing, denotes the noise signal after mixing, denotes the frequency component of the target scattering echo-jamming signal and direct wave-jamming signal respectively.
[0111] The mixed signal is low-pass filtered (LPF) with the radar i receiver bandwidth B i , and only when the jamming signal frequency band, jamming signal frequency band and radar i transmit signal frequency band have overlap, the jamming signal and jamming signal can enter the correlation processor of radar i.
[0112] In the embodiment, the step S2 is specifically as follows:
[0113] S21, analyze the conditions of radar frequency domain interference and mutual interference;
[0114] The frequency aiming probability of the interference on the m+1th pulse of the radar i is defined as:
[0115]
[0116] Where, ∩ represents the intersection operator, ΔB i and ΔB J respectively represent the intermediate frequency bandwidth of the receiver of the radar i and the instantaneous bandwidth of the interference, f J,m represents the center frequency of the interference. Similarly, the frequency aiming probability of the radar k transmission signal causing mutual interference on the m+1th pulse of the radar i is defined as:
[0117]
[0118]
[0119] Where, ΔB i and ΔB k respectively represent the intermediate frequency bandwidth of the receiver of the radar i and the instantaneous bandwidth of the radar k transmission signal; and respectively represent the frequency aiming probability of the scattering echo mutual interference signal and the direct wave mutual interference signal; represents the aiming probability of the radar k transmission signal on the main lobe of the radar i, and its expression is as follows:
[0120]
[0121] The interference signal and the mutual interference signal will cause interference on the radar i if and only if the frequency aiming probability of the radar i is not 0.
[0122] S22, establish a multi-radar frequency domain resource scheduling optimization model based on the SIJNR criterion;
[0123] The signal-interference-jamming-noise ratio (SIJNR) of the m+1th pulse of the radar i can be defined as:
[0124]
[0125] Where, P S,i,m , P J,i,m and P N,i respectively represent the actual power of the target echo signal received at the m+1th pulse of the radar i, the actual power of the interference signal, and the noise signal power, PI,k,i,m and P D,k,i,m respectively represent the actual power of the scattering echo mutual interference signal and the direct wave mutual interference signal received at the m+1th pulse of radar i; and P J,i,m , P I,k,i,m and P D,k,i,m satisfy:
[0126]
[0127]
[0128]
[0129] wherein γ D,k,i represents a complex amplitude parameter of the direct wave mutual interference signal transmitted by radar k and received by radar i.
[0130] According to the radar SIJNR model, the frequency domain resource scheduling model can be constructed as:
[0131]
[0132] wherein represents a set of selectable frequency points of radar i.
[0133] In the embodiment, the step S3 is specifically as follows:
[0134] S31, a generalized Markov decision process is established;
[0135] Each radar is regarded as an agent, and the multi-radar cooperative anti-jamming and anti-mutual interference process is constructed as a generalized Markov decision process, which can be represented by a five-tuple .
[0136] (1) Agent set The N frequency agile radars constitute the agent set;
[0137] (2) Action set represents the joint action space of the N agents, that is, the selectable frequency point range of all radars;
[0138] wherein a i,m = f i,m represents the action of radar i at the mth time step, a m = [f 1,m , f 2,m ,..., f N,m ] T represents the joint action of the agent at the mth time step (the m+1th pulse), (·) T represents a transpose operator.
[0139] (3) State set s represents all the optional jamming frequency points of the jammer, m J,m s represents the state presented by the jammer at the mth time step;
[0140] (4) State transition probability P: p(s'|s, a) represents the probability of the agent transitioning to state s' when performing action a in state s, and its definition is:
[0141] p(s'|s, a) = P{s m =s'|s m-1 =s, a m-1 =a} (18)
[0142] (5) Reward set represents the joint reward space of N agents, represents the reward set of agent i under policy π, in this embodiment, the SIJNR of each radar at the mth time step constitutes the reward set, i.e., r i,m (a i,m ,a -i,m )=SIJNR i,m ;
[0143] where a -i,m represents the joint action of other agents except agent i.
[0144] S32, establish a value function optimization model;
[0145] In the process of multi-radar cooperative anti-jamming and anti-interference, each radar is regarded as an agent, and the confrontation between the radar and the jammer can be regarded as the interaction process between the agent and the environment. Each CPI can be regarded as a scene, and each pulse can be regarded as each time step. Then the problem can be rewritten as:
[0146]
[0147] where, represents the expectation operator, γ ∈ [0, 1] represents the discount factor, π i :s→a i represents the policy of agent i.
[0148] The problem can be written as a value function optimization model as follows:
[0149]
[0150] where, Qi (s,a i ,a -i π represents the action value function of agent i. i (s,a i ) represents agent i in policy π i Next, action a i The probability distribution.
[0151] like Figure 3 and Figure 4 As shown in this embodiment, the specific solution process of the DQJIE algorithm in step S4 is as follows:
[0152] S41. Set up I coherent processing intervals CPI for training. Let j = 1, treat each CPI as a scene, initialize the discount factor γ ∈ [0,1], the learning rate η ∈ (0,1), and the state space. Action space Q-table for each agent Q i (s,a i ,a -i ) and the global Q-table Q(s,a), and begin the iteration for each scene;
[0153] S42. Let m = 0, treat each pulse in each CPI as a time step, and initialize s0 = f. J,0 ;
[0154] S43: Let i = 1, and start the iteration for each step of agent i;
[0155] S44. Select action a based on the ε-greedy policy for the (m+1)th impulse of agent i. i,m
[0156] S45, Action a based on agent i i,m and the current time interference carrier frequency s m =f J,m Receive reward r i,m And obtain the interference carrier frequency s at the next moment. m+1 =f J,m+1 ;
[0157] S46. Update the Q function Q of agent i. i (s,a i ,a -i ):
[0158]
[0159] Where s = s m a i =a i,m a -i =a-i,m , r i = r i,m , denotes the set of nodes having a communication link with agent i, denotes the number of elements in the set ;
[0160] S47, agent i shares its Q-table Q i (s, a i , a -i ) to its neighbor nodes, i.e., the nodes having a communication link with agent i;
[0161] S48, let i = i + 1, if i ≤ N, return to step S44, otherwise execute step S49;
[0162] S49, update the global Q-function Q(s, a):
[0163]
[0164] wherein s = s m , a i = a i,m , a -i = a -i,m .
[0165] S410, let m = m + 1, if m < M, return to step S43, otherwise execute step S411;
[0166] S411, let j = j + 1, if j ≤ I, return to step S42, otherwise execute step S412;
[0167] S412, output the Q-table Q i (s, a i , a -i ) of each agent and the global Q-table Q(s, a).
[0168] Each agent learns the frequency scanning strategy of the interference while interacting with the interference environment, thereby avoiding the frequency band where the interference is located, and in the process, interacts with other agents through the communication link, shares the action selection and Q-table at each moment, and then each agent corrects its own Q-table, forming a “transmission-feedback-correction” closed loop process, thereby realizing anti-interference in the frequency domain while effectively suppressing mutual interference.
[0169] The present application also provides another embodiment for simulating, verifying and analyzing the method of the present application:
[0170] In the process of networked radar cooperative detection, each radar realizes anti-jamming and mutual interference by inter-pulse frequency agility. Each radar transmits LFM signal, which is scattered by the target to form echo and is received by the radar, and then signal processing is performed.
[0171] The number of radar nodes in the networked radar is set to 3, the initial carrier frequency of the radar is 4 GHz, the signal bandwidth is 20 MHz, the signal pulse width is 20 μs, PRT = 400 μs, each radar transmits M = 100 pulses in one CPI, the coordinates of the radar in the Cartesian coordinate system are x R,1 = [0 km, 0 km, 0.3 km] T , x R,2 = [25 km, 30 km, 1.5 km] T and x R,3 = [-20 km, 35 km, 0.8 km] T , the coordinates of the target are x T,1 = [-120 km, 100 km, 10 km] T , the speed is v = 650 m / s, the target RCS is 10 m 2 , the instantaneous bandwidth of the jamming signal is 300 MHz, the jamming frequency sweeping mode is "random-step", that is, the frequency is randomly selected with a probability of 30%, and the frequency is stepped with a probability of 70%, and the jamming step frequency interval is 150 MHz. The frequency hopping range of the radar and the jamming is set to [4 GHz, 4.6 GHz]. The discount factor is γ = 0.8, the learning rate is η = e -0.01M , the ε in the greedy strategy is set to uniformly decrease from 0.3 to 0.15 at each time step, and the number of iteration CPI is 200.
[0172] In order to better show the effectiveness of the proposed DQJIE algorithm in resisting jamming and mutual interference, the following four methods are selected as the control: (1) frequency band division Q learning strategy (FBDivQL-FA), that is, each radar realizes Q learning algorithm in the frequency band divided by it, and the frequency hopping interval of each radar is 200 MHz; (2) independent Q learning strategy (IQL-FA), that is, all radars independently realize Q learning algorithm in a certain frequency band without considering mutual interference; (3) fixed carrier frequency strategy (Fix-CF), that is, the carrier frequencies of all radars are fixed; (4) random frequency selection strategy (Rand-FA), that is, all radars perform random frequency agility in a certain frequency band.
[0173] Figure 5 The comparison results of the total number of jamming and mutual interference pulses of each radar under different strategies are shown in the following figures, Figure 5 (a) is the comparison results of the total number of jamming and mutual interference pulses of radar 1 under different strategies, Figure 5 (b) is the comparison results of the total number of jamming and mutual interference pulses of radar 2 under different strategies,Figure 5 (c) is the comparison result of the total number of radar 3 pulses interfered by different strategies. As can be seen from the figure, in all strategies, due to the randomness of frequency selection, the effect of the random frequency selection strategy is relatively the worst, and in this strategy, the proportion of radar pulses interfered is about 65%. In the fixed carrier frequency strategy, there is no mutual interference, but the proportion of interfered pulses is still as high as 50%. The independent Q-learning strategy is relatively better than the previous two strategies in terms of anti-interference and mutual interference, but since this strategy does not consider mutual interference, it is defective in terms of comprehensive performance of anti-interference and anti-interference. Since there is no mutual interference in the frequency division Q-learning strategy, the anti-interference performance is suboptimal, and about 38% of the pulses are interfered. The DQJIE algorithm proposed in the method of the present application is the best in terms of comprehensive performance of anti-interference and anti-interference among all methods.
[0174] Figure 6 Fig. 4 is a comparison chart of the total number of pulses interfered by each radar under different strategies, Figure 6 (a) is a comparison chart of the total number of pulses interfered by radar 1 under different strategies, Figure 6 (b) is a comparison chart of the total number of pulses interfered by radar 2 under different strategies, Figure 6 (c) is a comparison chart of the total number of pulses interfered by radar 3 under different strategies. Since there is no mutual interference in the frequency division Q-learning strategy and the fixed carrier frequency strategy, Figure 6 only the comparison results of the other three methods are shown. In the random frequency selection strategy, about 20% of the pulses are interfered. Since in the independent Q-learning strategy, the reward of the agent is only affected by the interference environment, the number of pulses interfered is gradually increasing with iterations. The DQJIE algorithm proposed in the method of the present application comprehensively considers the influence of interference and mutual interference on the reward, has relatively optimal anti-interference performance, and effectively avoids mutual interference between radars.
[0175] Figure 7 Fig. 5 is a comparison chart of the SIJNR of each radar under different strategies, Figure 7 (a) is a comparison chart of the SIJNR of radar 1 under different strategies, Figure 7 (b) is a comparison chart of the SIJNR of radar 2 under different strategies, Figure 7 (c) is a comparison chart of the SIJNR of radar 3 under different strategies. In this embodiment, the SIJNR is calculated after pulse compression and non-coherent processing. As can be seen from the figure, the SIJNR of the DQJIE algorithm proposed in the method of the present application after pulse compression and non-coherent processing is obviously higher than that of other methods. Compared with the frequency division Q-learning strategy and the independent Q-learning strategy, the signal processing gain is increased by 3 dB and 4 dB respectively; compared with the random frequency selection strategy and the fixed carrier frequency strategy, the signal processing gain is increased by more than 5 dB.
[0176] To sum up, the method of the application constructs the anti-interference and anti-interference process of multiple radars into a generalized Markov decision process, takes the interference frequency band as a state, takes the radar frequency band as an action, takes the SIJNR as a reward function, and each radar intelligently selects the radar frequency band according to the interference frequency band at the current time, and the radars interact and feed back the action in real time through the communication link based on the graph theory model, and correct the respective Q table, so as to effectively suppress the mutual interference in the system while resisting the frequency sweeping interference, and effectively improve the detection performance of the networked radars.
[0177] Those skilled in the art will realize that the embodiments described herein are for the purpose of illustration and should not be construed as limiting the scope of the present application. The present application contemplates various modifications and changes that can be made thereto without departing from the spirit and scope of the present application. Any and all such modifications, variations, or equivalents that fall within the scope of the present application should, therefore, be considered within the scope of the claims that follow and any equivalents thereof.
Claims
1. A decentralized Q-learning-based networking radar jamming and mutual jamming elimination method, the specific steps being as follows: S1. According to the phased array radar signal processing flow, a target echo signal, a jamming signal and a mutual jamming signal model are established; S2. The conditions of radar being jammed and being mutually jammed are analyzed, and a multi-radar frequency domain resource scheduling optimization model based on the SIJNR criterion is established according to the definition of the frequency aiming probability; S21. The conditions of radar frequency domain being jammed and being mutually jammed are analyzed; Interfering with radar The The probability of frequency acquisition of the first pulse is defined as: (1) wherein denotes the radar The carrier frequency of the first transmitted pulse, denotes the intersection operator, and denotes the radar interferer's instantaneous bandwidth and the radar receiver's intermediate frequency bandwidth, respectively; and denotes the center frequency of the radar The frequency aiming probability of the first pulse to produce mutual interference is: (2) (3) wherein and respectively denote the mid-amp bandwidth of the radar receiver and the instantaneous bandwidth of the radar transmit signal; and respectively denote the frequency pointing probability of the scattered echo cross- signal and the direct wave cross-signal; denotes the pointing probability of the radar transmit signal to the main lobe of the radar receiver, which is expressed as follows: (4) Interference signals and mutual interference signals to radar will cause interference if and only if the frequency aiming probability for radar is not 0; S22. A multi-radar frequency domain resource scheduling optimization model based on the SIJNR criterion is established; The SIJNR definition expression for a radar pulse is given by (5) wherein , and denote the actual power of the target echo signal, the actual power of the jammer signal and the noise signal power, respectively, received at the radar first pulse, and denote the actual power of the scattered echo-jammer signal and the direct wave-jammer signal, respectively, received at the radar first pulse, , and , , denote complex parameters related to the propagation loss, denote the complex amplitude parameter of the radar received radar transmitted direct wave-jammer signal; According to the radar SIJNR model, the frequency domain resource scheduling model is constructed as: (6) wherein, represents a radar a set of optional frequencies, represents a radar a number of pulses transmitted within one coherent processing interval; S3. Each radar is regarded as an intelligent agent, the multi-radar cooperative anti-jamming and anti-mutual jamming process is constructed as a generalized Markov decision process, and the multi-radar frequency domain resource scheduling optimization model based on the SIJNR criterion in step S2 is converted into a value function optimization model; S4. The interference and mutual interference suppression algorithm based on decentralized Q-learning is used to solve the dynamic multi-agent frequency domain resource scheduling problem in step S2, and the frequency agility strategy of each radar is obtained.
2. The networking radar jamming and mutual jamming elimination method based on decentralized Q-learning according to claim 1, characterized in that, The step S1 is specifically as follows: Provided are The radars work in a time-synchronized manner to detect targets in cooperation, and the radars communicate in a time-synchronized manner. The linear frequency modulation signal is transmitted by each radar, and the carrier frequency of the signal transmitted by each radar is changed between pulses. The form of the transmitted signal of each radar is represented as (7) wherein, denotes a time instant, denotes a radar number of pulses transmitted within one coherent processing interval, denotes a pulse repetition time, denotes a radar the carrier frequency of the thtransmitted pulse, which can be denoted as , denotes a radar initial carrier frequency of the radar, denotes a frequency hopping interval, denotes a frequency encoding sequence; denotes a baseband chirp signal, ; Radar The transmitted signal probes the target. Assuming that both the jammer signal and the mutual interference signal enter through the side lobes, then the radar will receive the scattered echo signal from the target, the jammer signal and the mutual interference signal in the main lobe. The radar The received echo signal is represented as: (8) wherein the five terms on the right side of equation (2) represent, respectively, the target-scattered echo signal of the radar , the target-scattered mutual interference signal of the radar , the direct-wave mutual interference signal of the radar , the jamming signal and the noise signal; represents the two-way time delay caused by the distance between the radar and the target, represents the distance between the radar and the target, represents the radial velocity of the target relative to the radar , represents the speed of light; represents the time delay caused by the target-scattered signal of the radar transmitted to the radar , represents the velocity of the target relative to the radar in the radar transmitting and receiving paths, represents the one-way time delay caused by the distance between the radar and the radar , represents the Doppler shift of the target relative to the radar , represents the Doppler shift of the target relative to the radar in the radar transmitting and receiving paths; represents the signal wavelength of the radar , , , , and represent the complex parameters related to the propagation loss; the amplitudes of the target echo signal, the jamming signal, the scattered echo mutual interference signal and the direct-wave mutual interference signal are represented as: (9) (10) (11) (12) wherein , , and represent the peak transmit power, the transmit antenna gain, the receive antenna gain and the system loss of the radar ; represents the radar signal wavelength, represents the radar signal wavelength; represents the radar cross section of the target relative to the radar ; , and represent the peak transmit power, the transmit antenna gain and the wavelength of the jammer, respectively. The radar The received echo signal is down-converted to obtain the expression as follows: (13) wherein , , denotes the mixed interference signal, denotes the mixed noise signal, , denote the frequency components of the target-scattered echo interference signal and the direct-path wave interference signal, respectively; The mixed signal is filtered by a low-pass filter Receiver bandwidth The low-pass filter is used to limit the frequency band of the mixed signal, and the interference signal and the mutual interference signal will enter the correlation processor of the radar only when the frequency bands of the interference signal and the mutual interference signal overlap with the frequency band of the radar transmission signal.
3. The networking radar jamming and mutual jamming elimination method based on decentralized Q-learning according to claim 1, characterized in that, The step S3 is specifically as follows: S31. A generalized Markov decision process is established; Each radar is regarded as an agent, and the process of multi-radar cooperative anti-jamming and anti-interference is constructed as a generalized Markov decision process, which is represented by a five-tuple : (1) a set of agents : , the N transceivers constitute the set of agents; (2) Action set : represents the joint action space of N agents, i.e., the range of available frequency points for all radars; wherein, represents a radar At the action at the represents the joint action of the agent at the represents the joint action of the agent at the represents the transpose operator; (3) State set : represents the range of all possible jamming frequencies of the jammer, represents the state that the jammer presents at the th time step. (4) State transition probability : represents the probability that an agent performs action in state and transitions to state , which is defined as: (14) (5) reward set : represents the joint reward space of N agents, represents the reward set of agent under policy , and the SIJNR of each radar at the th time step constitutes the reward set, i.e. ; wherein, represents the joint action of other agents except the agent S32. A value function optimization model is established; Rewrite as: Rewrite as: (15) wherein, denotes the expected operator, denotes the discount factor, denotes the policy of the agent . Write the problem as a value function optimization model gives: (16) wherein, represents the action-value function of an agent , represents the probability distribution of an action under a policy of an agent .
4. The networking radar jamming and mutual jamming elimination method based on decentralized Q-learning according to claim 3, characterized in that, In the step S4, the DQJIE algorithm solving process is specifically as follows: S41, set co-training a phase tracking interval CPI, let each CPI be a scene, initialize discount factor learning rate state space action space Q table of each agent and global Q table start iteration of each scene; S42, let Initialize ; S43: Let , for each agent Start iteration of each step; S44, selecting an action for the agent based on the policy S45, based on the agent action and the current time interference carrier frequency get the reward , and get the next time interference carrier frequency ; S46, updating the agent Q-function of the agent : (17) wherein, , , , , denotes the set of nodes having a communication link with the agent , denotes the number of elements in the set . S47, agent Q-table for itself to its neighbor nodes, i.e. to the agents nodes having a communication link; S48, let , if , then return to step S44, otherwise execute step S49; S49, updating the global Q-function : (18) wherein , , ; S410, let , if , then return to step S43, otherwise execute step S411; S411, Order ,like If the condition is met, return to step S42; otherwise, proceed to step S412. S412、output the Q table of each agent and the global Q table .