Radar intelligent interference decision-making method, device and equipment and storage medium
By using the dual Q learning algorithm and the neural network structure of the GRU layer and the multi-head self-attention layer in the radar intelligent interference decision-making method, the optimal interference frequency is calculated, which solves the problem of combating frequency agile radar in the existing technology, and achieves efficient interference effect in complex dynamic environments.
Patent Information
- Application Number
- CN202311271364.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art is difficult to effectively combat frequency agile radars in complex dynamic environments, especially in the absence of radar prior information, and traditional interference decision-making methods are difficult to cope with the dynamic changes of frequency agile radars.
A radar intelligent interference decision-making method is adopted, and the trained target radar interference decision model is used to obtain and analyze the radar signal, and combine the dual Q learning algorithm and the neural network structure of the GRU layer and the multi-head self-attention layer to calculate the optimal interference frequency to counter frequency agile radar.
With large interference decision space and less radar prior information, the optimal interference frequency can be decided more quickly, effectively counter sub-pulse level frequency agile radar, and achieve a higher interference success rate.
Smart Images

Figure CN120233309A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of radar jamming, and particularly to a radar intelligent jamming decision-making method, device, equipment and storage medium. Background Art
[0002] Intelligent jamming decision-making is an important part of cognitive electronic warfare (CEW). With the continuous improvement of the radar's adaptability, higher requirements are put forward for more flexible jamming decision-making of jammers. Currently, frequency agile radars (FA Radars) have been widely used because they can effectively counter narrowband targeting active jamming and have advantages such as increasing the detection range, improving the angle measurement accuracy, and suppressing sea clutter.
[0003] However, since a frequency agile radar changes its carrier frequency within an extremely short time, while traditional jamming decisions usually interfere with specific frequencies and waveforms of the radar system, it is difficult for traditional jamming techniques to effectively jam the continuously changing operating frequencies of frequency agile radars. In addition, in actual application scenarios, the frequency hopping strategy of frequency agile radars is often unknown or even completely random, making it difficult to obtain the prior information of the radar, and even if obtained, there may be deviations. Traditional jamming decisions require a detailed model of the electromagnetic spectrum environment and the jammed party to solve the optimal jamming strategy under specific conditions, so it is difficult to cope with the dynamically changing electromagnetic environment. The currently proposed Q-learning algorithm that can adapt to dynamically changing environments and generate optimal jamming strategies, although achieving a large jamming-to-signal ratio (JNSR), has the disadvantage of poor convergence in solving problems with a large decision space and does not consider the jamming decision problem of sub-pulse-level frequency agile radars.
[0004] Therefore, how to overcome the deficiencies of the prior art and make jamming decisions on unknown frequency agile radars in a complex dynamic environment is an issue that still needs to be further solved in this field. Summary of the Invention
[0005] In view of this, the purpose of the present application is to provide a radar intelligent jamming decision-making method, device, equipment and storage medium, which are applied to cognitive jammers and can more quickly decide the optimal jamming frequency in the case of a large jamming decision space and less radar prior information, effectively counter sub-pulse-level frequency agile radars, and achieve a high jamming success rate. The specific solutions are as follows:
[0006] In a first aspect, the present application discloses a radar intelligent jamming decision-making method, which is applied to a cognitive jammer and includes:
[0007] Obtain the radar signal transmitted by the target frequency agile radar to be jammed, and analyze the radar signal to obtain the analyzed radar signal;
[0008] Input the analyzed radar signal into the trained target radar jamming decision model to calculate the optimal frequency for jamming the target frequency agile radar, and obtain the optimal jamming frequency; wherein, the target radar jamming decision model is a model obtained by training the initial radar jamming decision model constructed based on the duel DQN network including the target NNs network using a sample set; the sample set is the interactive experience data generated during the environmental interaction between the frequency agile radar and the cognitive jammer through the pre-created interactive environment model, the duel DQN network updates its parameters using the double Q-learning algorithm, and the target NNs network includes a GRU layer and a multi-head self-attention layer;
[0009] Adjust the frequency of the currently transmitted jamming signal to the optimal jamming frequency to jam the target frequency agile radar.
[0010] Optionally, before obtaining the radar signal transmitted by the target frequency agile radar to be jammed, it further includes:
[0011] Obtain the frequency change space of the target frequency agile radar and the frequency selection space of the cognitive jammer;
[0012] Judge whether the frequency selection space contains the frequency change space;
[0013] If the frequency selection space contains the frequency change space, execute the step of obtaining the radar signal transmitted by the target frequency agile radar to be jammed.
[0014] Optionally, the obtaining of the frequency change space of the target frequency agile radar includes:
[0015] Obtain all the available emission frequency bands and pulse repetition interval information of the target frequency agile radar through an electronic intelligence system to obtain the frequency change space of the target frequency agile radar.
[0016] Optionally, the radar intelligent jamming decision method further includes:
[0017] Conduct environmental interaction between the frequency agile radar and the cognitive jammer through the pre-created interactive environment model, and collect the interactive experience data generated during the interaction process;
[0018] Use the interactive experience data as a sample set, and screen out a preset number of mini-batch data sets from the sample set according to the weight size;
[0019] Train an initial radar jamming decision-making model constructed based on a dueling DQN network including a target NNs network using the small batch data set, and update the parameters of the dueling DQN network using the double Q-learning algorithm during the training process to obtain the target radar jamming decision-making model.
[0020] Optionally, the screening of a preset number of small batch data sets from the sample set according to the weight size includes:
[0021] Screen a preset number of small batch data sets from the sample set located in the prioritized experience replay buffer according to the weight size.
[0022] Optionally, the target NNs network further includes a noise linear network for processing the output of the multi-head self-attention layer and inputting the processing result into the state value function and the advantage function in the dueling DQN network.
[0023] Optionally, the radar intelligent jamming decision-making method further includes:
[0024] Evaluate the jamming effect of the optimal jamming frequency using the jamming-to-signal ratio.
[0025] In a second aspect, the present application discloses a radar intelligent jamming decision-making device applied to a cognitive jammer, including:
[0026] A radar signal acquisition module for acquiring radar signals emitted by a target frequency-agile radar to be jammed;
[0027] A radar signal analysis module for analyzing the radar signals to obtain the analyzed radar signals;
[0028] A jamming frequency calculation module for inputting the analyzed radar signals into the trained target radar jamming decision-making model to calculate the optimal frequency for jamming the target frequency-agile radar to obtain the optimal jamming frequency; wherein, the target radar jamming decision-making model is a model obtained by training an initial radar jamming decision-making model constructed based on a dueling DQN network including a target NNs network using a sample set; the sample set is interaction experience data generated during the process of environmental interaction between a frequency-agile radar and a cognitive jammer through a pre-created interaction environment model, the parameters of the dueling DQN network are updated using the double Q-learning algorithm, and the target NNs network includes a GRU layer and a multi-head self-attention layer;
[0029] A transmission frequency adjustment module for adjusting the frequency of the currently transmitted jamming signal to the optimal jamming frequency to jam the target frequency-agile radar.
[0030] In a third aspect, the present application discloses an electronic device, including a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the foregoing radar intelligent interference decision-making method is implemented.
[0031] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing radar intelligent interference decision-making method is implemented.
[0032] It can be seen that the present application is applied to a cognitive jammer. First, a radar signal transmitted by a target frequency-agile radar to be jammed is acquired, and the radar signal is analyzed to obtain the analyzed radar signal. Then, the analyzed radar signal is input into a trained target radar interference decision-making model to calculate the optimal frequency for jamming the target frequency-agile radar, and the optimal jamming frequency is obtained; wherein, the target radar interference decision-making model is a model obtained by training an initial radar interference decision-making model constructed based on a duel DQN network including a target NNs network using a sample set; the sample set is interactive experience data generated during the process of environmental interaction between a frequency-agile radar and a cognitive jammer through a pre-created interactive environment model. The duel DQN network uses a double Q-learning algorithm to update parameters. The target NNs network includes a GRU layer and a multi-head self-attention layer. Finally, the frequency of the currently transmitted jamming signal is adjusted to the optimal jamming frequency to jam the target frequency-agile radar. By adopting a duel DQN network using a double Q-learning algorithm and an NNs network including a GRU layer and a multi-head self-attention layer, the present application can more quickly decide the optimal jamming frequency in the case of a large interference decision space and less radar prior information. Compared with traditional reinforcement learning algorithms, it can effectively counter sub-pulse-level frequency-agile radars and achieve a better jamming effect. Description of the Drawings
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0034] Figure 1 It is a flowchart of a radar intelligent interference decision-making method disclosed in the present application;
[0035] Figure 2 It is a schematic diagram of the confrontation scenario between a cognitive jammer and a frequency-agile radar disclosed in the present application;
[0036] Figure 3Schematic diagram of a specific frequency agile radar pulse train disclosed in this application;
[0037] Figure 4 Schematic diagram of a specific frequency selection scheme for a frequency agile radar disclosed in this application;
[0038] Figure 5 Interaction block diagram of a radar and a cognitive jammer based on reinforcement learning disclosed in this application;
[0039] Figure 6 Schematic diagram of a specific intelligent jamming strategy algorithm block diagram disclosed in this application;
[0040] Figure 7 Schematic diagram of a specific Dueling DQN network structure disclosed in this application;
[0041] Figure 8 Schematic diagram of a specific NNs network structure disclosed in this application;
[0042] Figure 9 Schematic diagram of a specific state value network and advantage network structure disclosed in this application;
[0043] Figure 10 Schematic diagram for comparison of a radar intelligent jamming decision method disclosed in this application;
[0044] Figure 11 Schematic diagram for comparison of the benefits of a jammer with and without prior information disclosed in this application;
[0045] Figure 12 Schematic diagram of the structure of a radar intelligent jamming decision device disclosed in this application;
[0046] Figure 13 Schematic diagram of the structure of an electronic device disclosed in this application. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0048] The embodiments of this application disclose a radar intelligent jamming decision method, which is applied to a cognitive jammer. Refer to Figure 1 as shown, this method includes:
[0049] Step S11: Obtain the radar signal transmitted by the target frequency agile radar to be jammed, and analyze the radar signal to obtain the analyzed radar signal.
[0050] It should be noted that radar jamming decisions can be divided into two categories: active jamming and passive jamming. Compared with passive jamming, active jamming has greater potential in terms of jamming effectiveness. The radar intelligent jamming decision-making scheme proposed in this application belongs to active jamming, and the execution entity is a cognitive jammer (CJ). It is specifically applied to the scenario of the confrontation between a cognitive jammer and a frequency agile radar (i.e., FA Radar). The task of the cognitive jammer is to protect its own target object (Target) and reduce the probability of the target object being detected by the frequency agile radar. Specifically, for the confrontation scenario between the frequency agile radar and the cognitive jammer, see Figure 2 As shown, it includes a target object, a cognitive jammer, and a frequency agile radar. Additionally, it should be noted that in this embodiment, the cognitive jammer used for jamming decision-making is similar to the frequency agile radar, and both can perform frequency agility on sub-pulses, that is, they have the frequency agility function at the sub-pulse level, and can randomly select frequency bands between pulses, sense environmental changes, and adaptively adjust the jamming strategy; among them, the jamming strategy specifically refers to the frequency of the jamming signal transmitted by the cognitive jammer.
[0051] In this embodiment, when the cognitive jammer and the protected target object enter the flight space that can be jammed by the target frequency agile radar, first receive the radar signal transmitted by the above target frequency agile radar, and then perform an analysis operation on the received radar signal to obtain the analyzed radar signal; among them, the analyzed radar signal includes, but is not limited to, carrier information such as the frequency band range and bandwidth of the target frequency agile radar. Through this analyzed radar signal, the cognitive jammer can determine whether the jamming frequency is aligned with the radar frequency at the next moment.
[0052] In this embodiment, before obtaining the radar signal transmitted by the target frequency-agile radar to be jammed, it specifically further includes: obtaining the frequency change space of the target frequency-agile radar and the frequency selection space of the cognitive jammer; determining whether the frequency selection space contains the frequency change space; if the frequency selection space contains the frequency change space, then execute the step of obtaining the radar signal transmitted by the target frequency-agile radar to be jammed. It should be noted that the radar intelligent jamming decision-making scheme proposed in this application is implemented based on a pre-constructed interactive environment model of a frequency-agile radar, where the frequency-agile radar randomly changes its carrier frequency within a pulse, and moreover, the cognitive jammer has a frequency-agile function at the sub-pulse level. Since the interference to the target frequency-agile radar is achieved by transmitting a jamming frequency that changes, the frequency selection space of the cognitive jammer should contain the frequency change space of the target frequency-agile radar. Specifically, the frequency change space of the target frequency-agile radar can be obtained first, and then whether the above frequency selection space contains the above frequency change space. If the above frequency selection space contains the above frequency change space, it indicates that the cognitive jammer can perform frequency selection within the frequency band where the target frequency-agile radar is located, that is, a jamming decision can be made on the target frequency-agile radar.
[0053] Specifically, obtaining the frequency change space of the target frequency-agile radar may include: obtaining all the available emission frequency bands and pulse repetition interval information of the target frequency-agile radar through an electronic intelligence system to obtain the frequency change space of the target frequency-agile radar. In this embodiment, all the available emission frequency bands and pulse repetition interval information (PRI, Pulse Repetition Interval) of the target frequency-agile radar can be obtained through an electronic intelligence system (ELINT, Electronic Intelligence), and then the above frequency change space of the target frequency-agile radar can be obtained.
[0054] Specifically, referring to Figure 3 shown Figure 3 shows a pulse train transmitted by a specific sub-pulse agile frequency-agile radar. Figure 3 The frequency band range of the frequency-agile radar in L is F = [f H , the bandwidth is B = f H - f L . The sub-band is obtained by equally dividing the bandwidth B into M bands; each band range is F = [f L + d n Δf, f L + (d n + 1)△f], where d n is the frequency modulation code and is less than random integers with a frequency change step of Figure 3 The pulse train in includes N pulses with n = 1, 2, ..., N. The pulse repetition period is T. Each pulse is divided into K sub-pulses. The frequencies of the sub-pulses are pairwise different. All K sub-pulses together occupy the sub-frequency band where the sub-pulse is located. And the carrier frequency adopted by the k-th sub-pulse of the n-th pulse is f k (n) , further, refer to Figure 4 as shown in Figure 4 The frequency selection scheme that each sub-pulse may adopt when K = 4. At this time, the number of frequencies included in the interference frequency selection space of the cognitive jammer is K×M.
[0055] Step S12: Input the parsed radar signal into the trained target radar interference decision model to calculate the optimal frequency for interfering with the target frequency agile radar, and obtain the optimal interference frequency; wherein, the target radar interference decision model is a model obtained by training the initial radar interference decision model based on the dueling DQN network including the target NNs network using a sample set; the sample set is the interactive experience data generated during the environmental interaction between the frequency agile radar and the cognitive jammer through a pre-created interactive environment model. The dueling DQN network updates its parameters using the double Q-learning algorithm. The target NNs network includes a GRU layer and a multi-head self-attention layer.
[0056] In this embodiment, after parsing the radar signal to obtain the parsed radar signal, further, the parsed radar signal is used as the environmental state and input into the trained target radar interference decision model, so as to calculate the optimal frequency for interfering with the target frequency agile radar through the target radar interference decision model, that is, the optimal interference frequency. Among them, the target radar interference decision model is a model obtained by training the initial radar interference decision model based on the dueling DQN (i.e., Dueling Deep Q-network) network including the target NNs (Neural Networks) network using a sample set; the sample set is the interactive experience data generated during the environmental interaction between the frequency agile radar and the cognitive jammer through a pre-created interactive environment model; it should be noted that the dueling DQN network uses the double Q (Q-Learning) learning algorithm to update the parameters in the model, and the target NNs network specifically includes a GRU (Gated Recurrent Unit) layer and a multi-head self-attention layer (MultiHead-Self-Attention Layer).
[0057] Specifically, the training process of the target radar jamming decision-making model may include: performing environmental interaction on a frequency agile radar and a cognitive jammer through a pre-created interaction environment model, and collecting interaction experience data generated during the interaction process; using the interaction experience data as a sample set, and screening out a preset number of mini-batch data sets from the sample set according to the weight size; training an initial radar jamming decision-making model constructed based on a duel DQN network including a target NNs network using the mini-batch data set, and updating the parameters of the duel DQN network using a double Q-learning algorithm during the training process to obtain the target radar jamming decision-making model. In this embodiment, it is necessary to pre-create an interaction environment model for performing environmental interaction on a frequency agile radar and a cognitive jammer, and collecting interaction experience data generated during the interaction process, then using the collected interaction experience data as a sample set, and then screening out a preset number of mini-batch data sets from the above sample set according to the weight size, and then using the above mini-batch data set to train an initial radar jamming decision-making model constructed based on a duel DQN network including a target NNs network. Among them, the cognitive jammer performs interaction in the created above interaction environment model, learns the optimal jamming frequency for jamming, and adjusts the frequency of the transmitted jamming signal to maximize its long-term cumulative reward in the interaction environment model; in addition, during the training process, the parameters of the duel DQN network are updated using a double Q-learning algorithm to obtain the target radar jamming decision-making model. It should be noted that during the learning stage of the duel DQN network (that is, the process in which the cognitive jammer learns the radar frequency change law through interaction), the cognitive jammer agent can screen out a preset number of mini-batch data sets including the historical radar emission information of the target frequency agile radar from the prioritized experience replay buffer (PER, Prioritized Experience Replay Buffer), and then use the obtained mini-batch data set (mini batch) to train an initial radar jamming decision-making model constructed based on a duel DQN network including a target NNs network to realize the update of the network.
[0058] In a specific implementation manner, the screening out a preset number of mini-batch data sets from the sample set according to the weight size may specifically include: screening out a preset number of mini-batch data sets from the sample set located in the prioritized experience replay buffer according to the weight size. For example, screening out a preset number of mini-batch data sets from all the sample sets located in the PER in the order of decreasing weight.
[0059] It should be noted that the target radar interference decision-making model for calculating the optimal interference frequency in this application is implemented based on the deep reinforcement learning algorithm, but the traditional deep reinforcement learning algorithm has been improved. The specific improvements include:
[0060] Improved based on Dueling DQN by adding a GRU layer and an attention mechanism to address the deficiencies of traditional DQN. The introduction of the GRU layer allows the model to retain state information between multiple time steps, enhancing the model's stability against randomness and uncertainty in the environment. The attention mechanism helps the model automatically filter information more relevant to the current state, reducing interference from irrelevant information. In addition, thanks to the method of decomposing the Q-value function into a state-value function and an action-advantage function in Dueling DQN, the understanding of state value and action importance is improved, thus more accurately estimating the Q-value of each action. In an environment with completely random states, the above method can enable the model to converge to the optimal solution faster. It can be understood that the traditional deep reinforcement learning algorithm enables the agent to interact with the environment, allowing it to learn through trial and error to obtain the maximum cumulative reward, thereby achieving task optimization. As the basis of reinforcement learning, the Markov decision process (MDP, Markov Decision Process) is a mathematical model used to describe sequential decision-making problems. In the Markov decision process, the agent interacts with the environment, makes decisions based on the observed current state, executes actions, obtains rewards, and enters a new state. The immediate reward obtained is related to both the state and the action, and this process is continuously iterated until it ends. Specifically, the Markov decision process can be labeled as <S, A, P, R>; where S is the set of environmental states that the agent can observe, A is the set of actions that the agent can take in each state, P is the transition probability P a (s,s') = P(s t+1 = s′|s t = s,a t = a), representing the probability that the next state transfers to s′ after executing action a in state s, and R is the reward function representing the expected value of the immediate reward that the agent can obtain after executing action a in state s. Moreover, in the Markov decision process, the goal of the agent is to maximize the long-term cumulative reward. In this embodiment, the cognitive jammer is used as the agent in the deep reinforcement learning algorithm, and the target frequency-agile radar is used as the dynamically changing environment. The action space and the state space are both K×M optional pulse signal frequencies. Specifically, see Figure 5 as shown Figure 5Shows the interaction process between the target frequency agile radar and the cognitive jammer based on the Markov decision process, modeled according to the target frequency agile radar and the cognitive jammer, where the environmental state is S t ={f t (r) The total number of states is K×M, and the action of the agent is A t ={f t (j)}, f t (j) is the frequency of the pulse signal emitted by the cognitive jammer. The total number of actions is the same as the total number of states. The reward r t is the interference-to-signal ratio received by the target frequency agile radar, and its calculation formula is:
[0061]
[0062] where N is the number of radar sub-pulses with the same frequency as the interference sub-pulses, T is the pulse repetition interval. The goal of the cognitive jammer is to find an interference frequency selection strategy π * (a|s).
[0063] In addition, it should be noted that the Dueling DQN network in the target radar interference decision model is constructed based on the DQN deep reinforcement learning algorithm, but the DQN deep reinforcement learning algorithm is improved. It can be understood that, as a deep reinforcement learning algorithm based on Q-learning, the DQN deep reinforcement learning algorithm uses a deep neural network to approximate the Q-value function and stabilizes the training process through techniques such as Experience Replay and Target Network. During the experience replay process, the agent stores the state-action-reward sequences it has experienced in the replay buffer and randomly samples from it for training. However, it should be noted that Q-learning uses its own estimate to update the Q-value, which will lead to the propagation of bias, and maximization will cause the estimated target to be higher than the true value, that is, the DQN network trained by Q-learning has the problem of overestimation, and this overestimation is usually non-uniform. To solve the overestimation problem caused by the above maximization, this application uses the double Q-learning algorithm to update the parameters of the Dueling DQN network. See Figure 6 shown. The double Q-learning algorithm includes a frequency selection policy network and a frequency selection target network. The execution process of this algorithm is: first, use the frequency selection policy network to select an action based on the state st+1 to maximize the output of the DQN, and then calculate the value Q(s t+1, a). It should be noted that both the frequency selection policy network and the frequency selection target network adopt the Dueling DQN network. Among them, the expression of the frequency selection policy network Q policy (s, a, w) is:
[0064] q t = Q policy (o t , h t-1 , w);
[0065] In the formula, w is the network parameter of Q policy , which is generated by random initialization; o t is the observation of the cognitive jammer at time step t (i.e., the parsed radar information), and h t-1 is the memory vector of the cognitive jammer for time step t - 1 and before. It should be noted that when initializing time step t = 1, the memory vector h policy of Q t-1 is all 0.
[0066] Specifically, the expression of the frequency selection target network Q target (s, a, w - ) is:
[0067] q t ' = Q taget (o t , g t-1 , w - );
[0068] Among them, Q target and Q policy have the same structure; q t ' is the evaluation vector of the cognitive jammer for each selectable frequency at time step t; g t-1 is the memory vector of the cognitive jammer for time step t - 1 and before, that is, the memory of the frequency agile radar for time step t and before; w - is the network parameter of Q target .
[0069] Specifically, when using the Dueling DQN network to make interference decisions on the target frequency agile radar, the cognitive jammer inputs the observation information ot, such as the parsed radar information, into the current frequency selection policy network Q policy (o t , h t-1 , w), and then obtains the q-value vector q policy (o t , h t-1 , w) output by q t = [q t 1 , qt 2 , …, q t K×M , where q t i , i = 1, 2, ..., K×M, representing the interference effectiveness evaluation that the cognitive jammer will obtain when adopting frequency i at time step t. The formula for the cognitive jammer to select the interference pulse frequency at time step t is:
[0070]
[0071] where ε is the exploration rate that decays exponentially.
[0072] It should be noted that the Dueling DQN network in this application adds a target NNs network including a GRU layer and a multi-head self-attention layer to the traditional DQN network to solve the intelligent interference decision-making problem of modeling the cognitive jammer and the target frequency-agile radar as an MDP. In the Dueling DQN network, the goal of the cognitive jammer (i.e., the agent) is to find an interference frequency selection strategy π * (a|s). And, the target NNs network specifically further includes a Noisy Linear Network (NoisyNet) for processing the output of the multi-head self-attention layer and inputting the processing result into the state value function and the advantage function in the duel DQN network. See Figure 7 shown, this application decomposes the action value function Q π (s, a) into a state value function V π (s) and an advantage function A(s, a); where V π (s) is used to predict the expected return of a state, A(s, a) is used to calculate the advantage value of each action relative to the expected return. Further, the results of V π (s) and A(s, a) are combined to calculate the Q value of each action. By decomposing the action value function, the agent can learn the relative advantages of each action instead of learning the value of each action alone. Therefore, when the value differences of some actions of the agent in some states are not large, the advantages and disadvantages of each action can still be accurately evaluated. The specific calculation formula is:
[0073]
[0074]
[0075]
[0076] In this embodiment, see Figure 8 andFigure 9 As shown, a target NNs network including a GRU layer and a multi-head self-attention layer is added to the traditional DQN network. Specifically, at each time step (i.e., each interaction between the cognitive jammer and the target frequency-agile radar), the received radar pulse signal is first input into the GRU layer in the target NNs network, enabling the cognitive jammer to comprehensively consider the current observation o t (i.e., the radar pulse signal received in the current moment interaction) and the historical observations (i.e., the radar pulse signal received at the current moment) to switch the transmission frequency, thus solving the problem of insufficient single observation of the cognitive jammer; then, a multi-head (e.g., 8 heads) self-attention layer is used to learn the correlation between different features of the input, so as to better extract the relevant features and representations of the input data; finally, a linear layer processes the output of the multi-head self-attention layer to further transform it into a feature representation with higher dimension and richer semantic information, and the processed result is input into two identical linear layers as the input of the state value function and the advantage function. In this way, the model can more effectively learn the key information in the input sequence, that is, learn the key features in the input radar pulse signal.
[0077] In a specific embodiment, refer to Figure 9 As shown, the linear layer can use a noisy linear network to increase exploration by adding Gaussian noise to the weights and biases of the neural network, thereby improving the learning effect of the Dueling DQN network. Then, the output after passing through the noisy linear network is input to Figure 7 the state value function V π (s) and the advantage function A(s, a), and then the Q π (s, a; w), that is, the Q value of each action, is obtained.
[0078] Step S13: Adjust the frequency of the currently transmitted interference signal to the optimal interference frequency to interfere with the target frequency-agile radar.
[0079] In this embodiment, after calculating the optimal interference frequency using the target radar interference decision model, the cognitive jammer can adjust the frequency of the currently transmitted interference signal according to the above optimal interference frequency, thereby interfering with the above target frequency-agile radar and achieving the protection of the target object.
[0080] In this embodiment, after obtaining the optimal interference frequency, the interference effect of the optimal interference frequency can be evaluated using the jammer-to-signal ratio. In this embodiment, the jammer learns the radar frequency agility strategy by continuously interacting with the target frequency-agile radar through a target radar interference decision model constructed based on a duel DQN network including the target NNs network, and determines the optimal interference frequency according to the learned radar frequency agility strategy. During the process of interacting and learning the radar frequency change rule, the cognitive jammer can determine the optimal interference frequency by increasing the jammer-to-signal ratio (i.e., Jammer-to-Signal Ratio, JSR), that is, the objective function of the jammer is JSR. Specifically, the calculation formula for the jammer-to-signal ratio JSR received by the target frequency-agile radar is:
[0081]
[0082] Where,
[0083] In the formula, h t is the channel gain from the target frequency-agile radar to the target object, h j is the channel gain from the cognitive jammer to the target frequency-agile radar, L is the sidelobe loss of the cognitive jammer entering the target frequency-agile radar, σ is the RCS (Radar Cross Section), τ (R) is the pulse width of the target frequency-agile radar, A R is the pulse amplitude of the target frequency-agile radar, T is the pulse repetition interval (PRI, Pulse Repetition Interval), P R is the transmit power of the target frequency-agile radar, the pulse width of the cognitive jammer is τ (J) , the pulse amplitude is A J , P J is the transmit power of the cognitive jammer. It should be noted that the number of frequencies specifically included in the interference frequency space that the cognitive jammer can select is K×M.
[0084] Specifically, to verify the effectiveness of the optimal interference frequency, the Dueling DQN network proposed in this application can be compared with the DQN network and the Q-learning algorithm through simulation experiments. In a specific experimental method, the parameters of the DQN network and the Dueling DQN network structure are shown in Table 1. The only difference between the two network structure parameters is that the GRU layer and the multi-head attention layer are replaced with a linear layer. For the remaining parameters such as the exploration rate and the discount factor, the three methods are exactly the same. Moreover, the exploration rate decreases from 0.995 to 0.005 in the form of exponential decay, and the exploration rate decay rate remains consistent. Both the DQN network and the Dueling DQN network adopt the cosine learning rate warm-up technique, and the maximum learning rate is 0.009.
[0085] Table 1
[0086]
[0087] Specifically, the carrier frequency range of the frequency agile radar of the three methods is F = [10 GHz, 11 GHz], and it is divided into M = 10 sub-bands at equal intervals of 100 MHz as the frequency band range of the pulse. In addition, the pulse width τ is 40 ns, and each pulse has K = 4 sub-pulses. Therefore, the bandwidth of the sub-pulse is Δf = 25 MHz. In order to make each pulse fully utilize its bandwidth, the frequencies of the 4 sub-pulses are all different from each other in pairs. Therefore, there are a total of 24 combinations, that is, the state space of the frequency agile radar is 10×24 = 2240, and the action space and state space of the jammer are the same. In addition, the parameters of the frequency agile radar and the cognitive jammer in the three methods are shown in Table 2.
[0088] Table 2
[0089]
[0090] The Dueling DQN network, the DQN network and the Q-learning algorithm are used, and interference decisions are made according to the above parameters. Among them, the total number of pulse trains (episodes) is 100, and each episode has 10,000 time steps. The comparison of the total rewards of each episode is shown in Figure 10 as shown, through Figure 10It can be seen that the total reward of each episode of the above three methods increases with the increase of the number of interaction episodes, indicating that these three methods can enable the jammer to obtain the frequency agility law of the radar through continuous learning, thereby improving the jamming effect. Therefore, the effectiveness of the reinforcement learning algorithm for intelligent jamming decision-making is proved. In addition, by comparing the three methods, it can be seen that the Dueling DQN network proposed in this application can converge to the highest total reward relatively quickly. Therefore, compared with the other two methods, this application can further improve the learning rate and effect of the jammer's intelligent decision-making, has significant advantages, can converge to a stable state more quickly, and can converge to the global optimal solution.
[0091] Furthermore, the situation where the cognitive jammer can obtain some prior information of the frequency agile radar is verified. Specifically, the interference decision-making is verified for the situation where the cognitive jammer obtains the carrier frequency of the first sub-pulse of each radar pulse through interception and parameter estimation, and it is known that although the radar performs frequency agility between sub-pulses, it jumps within the sub-band range. The verification results of the above three methods are shown in Figure 11 as follows. Through Figure 11 It can be seen that when the frequency of the reconnaissance sub-pulse is used as the prior information, the above three methods can all improve the interference reward, because the acquisition of prior information can greatly reduce the interference decision-making space. After using the prior information, the interference success probability of the last episode is shown in Table 3. It can be seen from Table 3 that the interference success rates of the Q-learning algorithm and the DQN network have been significantly improved; among them, the interference success rate of the Q-learning algorithm has increased from 45.32% to 75.01%, and the DQN network has increased from 72.98% to 82.56%. The Dueling DQN network proposed in this application can achieve an interference pulse hit rate of more than 97% without radar prior information. When there is radar prior information, the interference hit rate is further improved to 99.41%, with a difference of less than 3%. Therefore, the Dueling DQN network proposed in this application has significant superiority compared with the Q-learning algorithm and the DQN network, and can achieve a very high interference success rate without the prior information of the radar sub-pulse.
[0092] Table 3
[0093]
[0094] As can be seen from the above, the Dueling DQN network proposed in this application decomposes the action value function into a state value function and an advantage function, which can more accurately estimate the value of the state and the advantages and disadvantages of each action, that is, can more accurately estimate the Q-value of each action, enabling the model to more effectively learn the key information in the input radar sequence. In an environment where the state is completely random, the model can converge to the optimal solution faster. Compared with traditional reinforcement learning algorithms, it can converge to a stable state more quickly and achieve a better interference effect in the case of a large interference decision space. By introducing the GRU layer, certain state information can be retained between multiple time steps, making the model more stable to the randomness and uncertainty in the environment. The multi-head self-attention mechanism can help the model automatically filter out information relevant to the current task, reducing the interference of irrelevant information. As can be seen from the above, the cognitive jammer in the embodiment of this application can perform frequency agility at the sub-pulse level and uses a radar interference decision model constructed by a dueling DQN network including a target NNs network to continuously interact with the frequency-agile radar to learn the radar frequency agility strategy, thereby obtaining the optimal interference frequency. Among them, the target NNs network includes a GRU layer and a multi-head self-attention layer, enabling the model to more stably learn the key information in the input radar sequence, thus solving the deficiencies of the DQN network. The GRU layer can synthesize current and historical observation information to switch the transmission frequency, solving the problem of insufficient single observation of the radar. At the same time, by introducing the attention mechanism, information relevant to the current task can be filtered out, reducing the interference of irrelevant information. Compared with traditional reinforcement learning algorithms, it can converge to a stable state more quickly and achieve a better interference effect in the case of a large interference decision space.
[0095] It can be seen that the embodiment of the present application is applied to a cognitive jammer. First, the radar signal transmitted by the target frequency agile radar to be jammed is acquired, and the radar signal is analyzed to obtain the analyzed radar signal. Then, the analyzed radar signal is input into the trained target radar jamming decision model to calculate the optimal frequency for jamming the target frequency agile radar, and the optimal jamming frequency is obtained. Among them, the target radar jamming decision model is a model obtained by training the initial radar jamming decision model based on the duel DQN network including the target NNs network using a sample set. The sample set is the interactive experience data generated during the environmental interaction between the frequency agile radar and the cognitive jammer through the pre-created interactive environment model. The duel DQN network uses the double Q-learning algorithm to update parameters. The target NNs network includes a GRU layer and a multi-head self-attention layer. Finally, the frequency of the currently transmitted jamming signal is adjusted to the optimal jamming frequency to jam the target frequency agile radar. By adopting the duel DQN network using the double Q-learning algorithm and the NNs network including the GRU layer and the multi-head self-attention layer, the present application can more quickly decide the optimal jamming frequency in the case of a large jamming decision space and less radar prior information. Compared with the traditional reinforcement learning algorithm, it can effectively counter the sub-pulse level frequency agile radar and achieve a better jamming effect.
[0096] Correspondingly, the embodiment of the present application also discloses a radar intelligent jamming decision device, which is applied to a cognitive jammer. Refer to Figure 12 As shown, the device includes:
[0097] A radar signal acquisition module 11, configured to acquire the radar signal transmitted by the target frequency agile radar to be jammed;
[0098] A radar signal analysis module 12, configured to analyze the radar signal to obtain the analyzed radar signal;
[0099] A jamming frequency calculation module 13, configured to input the analyzed radar signal into the trained target radar jamming decision model to calculate the optimal frequency for jamming the target frequency agile radar, and obtain the optimal jamming frequency. Among them, the target radar jamming decision model is a model obtained by training the initial radar jamming decision model based on the duel DQN network including the target NNs network using a sample set. The sample set is the interactive experience data generated during the environmental interaction between the frequency agile radar and the cognitive jammer through the pre-created interactive environment model. The duel DQN network uses the double Q-learning algorithm to update parameters. The target NNs network includes a GRU layer and a multi-head self-attention layer;
[0100] The transmission frequency adjustment module 14 is used to adjust the frequency of the currently transmitted interference signal to the optimal interference frequency to interfere with the target frequency agile radar.
[0101] Among them, for the specific working processes of the above-mentioned various modules, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.
[0102] It can be seen that the embodiment of the present application is applied to a cognitive jammer. First, a radar signal transmitted by a target frequency agile radar to be jammed is acquired, and the radar signal is analyzed to obtain an analyzed radar signal. Then, the analyzed radar signal is input into a trained target radar interference decision-making model to calculate the optimal frequency for interfering with the target frequency agile radar, and the optimal interference frequency is obtained. Among them, the target radar interference decision-making model is a model obtained by training an initial radar interference decision-making model constructed based on a duel DQN network including a target NNs network using a sample set. The sample set is interactive experience data generated during the process of environmental interaction between a frequency agile radar and a cognitive jammer through a pre-created interactive environment model. The duel DQN network updates its parameters using a double Q-learning algorithm. The target NNs network includes a GRU layer and a multi-head self-attention layer. Finally, the frequency of the currently transmitted interference signal is adjusted to the optimal interference frequency to interfere with the target frequency agile radar. By adopting a duel DQN network using a double Q-learning algorithm and an NNs network including a GRU layer and a multi-head self-attention layer, the present application can more quickly determine the optimal interference frequency in a large interference decision space with less radar prior information. Compared with traditional reinforcement learning algorithms, it can effectively counter sub-pulse-level frequency agile radars and achieve a better interference effect.
[0103] In some specific embodiments, before the radar signal acquisition module 11, it may further include:
[0104] A space acquisition unit for acquiring the frequency change space of the target frequency agile radar and the frequency selection space of the cognitive jammer;
[0105] A judgment unit for judging whether the frequency selection space contains the frequency change space;
[0106] An execution unit for, if the frequency selection space contains the frequency change space, executing the step of acquiring the radar signal transmitted by the target frequency agile radar to be jammed.
[0107] In some specific embodiments, the space acquisition unit may specifically include:
[0108] A frequency change space acquisition unit, configured to obtain all the available emission frequency bands and pulse repetition interval information of the target frequency agile radar through an electronic intelligence system, so as to obtain the frequency change space of the target frequency agile radar.
[0109] In some specific embodiments, the radar intelligent interference decision-making device may further include:
[0110] An environment interaction unit, configured to perform environment interaction on the frequency agile radar and the cognitive jammer through a pre-created interaction environment model;
[0111] A data acquisition unit, configured to acquire the interaction experience data generated during the interaction process;
[0112] A first data set acquisition unit, configured to use the interaction experience data as a sample set, and screen out a preset number of mini-batch data sets from the sample set according to the weight size;
[0113] A model training unit, configured to train an initial radar interference decision-making model constructed based on a duel DQN network including a target NNs network by using the mini-batch data set, and update the parameters of the duel DQN network by using a double Q learning algorithm during the training process to obtain the target radar interference decision-making model.
[0114] In some specific embodiments, the first data set acquisition unit may specifically include:
[0115] A second data set acquisition unit, configured to screen out a preset number of mini-batch data sets from the sample set in the prioritized experience replay buffer according to the weight size.
[0116] In some specific embodiments, the target NNs network further includes a noise linear network, configured to process the output of the multi-head self-attention layer and input the processing result into the state value function and the advantage function in the duel DQN network.
[0117] In some specific embodiments, the radar intelligent interference decision-making device may further include:
[0118] An evaluation unit, configured to evaluate the interference effect of the optimal interference frequency by using the interference-to-signal ratio.
[0119] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 13 which is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be regarded as any limitation to the scope of use of the present application.
[0120] Figure 13Schematic diagram of the structure of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the radar intelligent interference decision-making method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0121] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitations are made here.
[0122] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.
[0123] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the radar intelligent interference decision-making method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.
[0124] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the radar intelligent interference decision-making method disclosed above is implemented. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details are not repeated here.
[0125] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0126] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0127] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0128] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0129] The above has introduced in detail a radar intelligent interference decision method, device, equipment and storage medium provided by this application. Specific examples are used in this document to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A radar intelligent interference decision-making method, characterized in that, Applied to a cognitive jammer, including: Obtain the radar signal transmitted by the target frequency-agile radar to be jammed, and analyze the radar signal to obtain the analyzed radar signal; Input the analyzed radar signal into the trained target radar jamming decision-making model to calculate the optimal frequency for jamming the target frequency-agile radar, and obtain the optimal jamming frequency; wherein, the target radar jamming decision-making model is a model obtained by training the initial radar jamming decision-making model constructed based on the duel DQN network including the target NNs network using a sample set; the sample set is the interactive experience data generated during the environmental interaction between the frequency-agile radar and the cognitive jammer through a pre-created interactive environment model, the duel DQN network updates its parameters using the double Q-learning algorithm, and the target NNs network includes a GRU layer and a multi-head self-attention layer; Adjust the frequency of the currently transmitted jamming signal to the optimal jamming frequency to jam the target frequency-agile radar.
2. The radar intelligent interference decision-making method according to claim 1, wherein Before obtaining the radar signal transmitted by the target frequency-agile radar to be jammed, it further includes: Obtain the frequency change space of the target frequency-agile radar and the frequency selection space of the cognitive jammer; Judge whether the frequency selection space contains the frequency change space; If the frequency selection space contains the frequency change space, execute the step of obtaining the radar signal transmitted by the target frequency-agile radar to be jammed.
3. The radar intelligent interference decision-making method according to claim 2, wherein The obtaining of the frequency change space of the target frequency-agile radar includes: Obtain all the available emission frequency bands and pulse repetition interval information of the target frequency-agile radar through an electronic intelligence system to obtain the frequency change space of the target frequency-agile radar.
4. The radar intelligent interference decision-making method according to claim 1, wherein It further includes: Conduct environmental interaction between the frequency-agile radar and the cognitive jammer through a pre-created interactive environment model, and collect the interactive experience data generated during the interaction process; Use the interactive experience data as a sample set, and screen out a preset number of mini-batch data sets from the sample set according to the weight size; Use the mini-batch data set to train the initial radar jamming decision-making model constructed based on the duel DQN network including the target NNs network, and update the parameters of the duel DQN network using the double Q-learning algorithm during the training process to obtain the target radar jamming decision-making model.
5. The radar intelligent interference decision-making method according to claim 4, wherein The screening out of a preset number of mini-batch data sets from the sample set according to the weight size includes: Screen out a preset number of mini-batch data sets from the sample set located in the prioritized experience replay buffer according to the weight size.
6. The radar intelligent interference decision-making method according to claim 1, characterized in that, The target NNs network further includes a noise linear network, which is used to process the output of the multi-head self-attention layer and input the processing result into the state value function and the advantage function in the duel DQN network.
7. The radar intelligent interference decision-making method according to any one of claims 1 to 6, characterized in that It further includes: Evaluate the jamming effect of the optimal jamming frequency using the jamming-to-signal ratio.
8. A radar intelligent interference decision-making device, characterized in that, Applied to a cognitive jammer, including: A radar signal acquisition module, which is used to obtain the radar signal transmitted by the target frequency-agile radar to be jammed; A radar signal analysis module, which is used to analyze the radar signal to obtain the analyzed radar signal; An interference frequency calculation module, which is used to input the analyzed radar signal into the trained target radar interference decision model to calculate the optimal frequency for interfering with the target frequency agile radar, and obtain the optimal interference frequency; wherein, the target radar interference decision model is a model obtained by training the initial radar interference decision model constructed based on the confrontation DQN network including the target NNs network by using a sample set; the sample set is the interaction experience data generated during the environmental interaction of the frequency agile radar and the cognitive jammer through the pre-created interaction environment model, the confrontation DQN network updates parameters by using the double Q learning algorithm, and the target NNs network includes a GRU layer and a multi-head self-attention layer; A transmission frequency adjustment module, which is used to adjust the frequency of the currently transmitted interference signal to the optimal interference frequency to interfere with the target frequency agile radar.
9. An electronic device, characterized in that, It includes a processor and a memory; wherein, when the processor executes the computer program stored in the memory, it implements the radar intelligent interference decision method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It is used to store a computer program; wherein, when the computer program is executed by the processor, it implements the radar intelligent interference decision method according to any one of claims 1 to 7.