A radar intelligent anti-jamming method combining attention mechanism and A2C algorithm

By combining the attention mechanism and the A2C algorithm, an MDP model was constructed and reinforcement learning was performed. This solved the problems of high variance in policy gradient and insufficient signal feature extraction in radar anti-jamming technology, realized a high-precision frequency agility strategy, and improved the anti-jamming capability of the radar system in complex electromagnetic environments.

CN120669211BActive Publication Date: 2026-08-25XIAN LEITONG SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510923668.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2026-08-25
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

Existing radar anti-jamming technologies are insufficient to cope with the dynamic game of intelligent jammers. Policy gradient algorithms suffer from high variance, value function-based methods are inefficient in continuous action spaces, and radar signal feature extraction lacks the ability to focus on key jamming signals, resulting in insufficient decision robustness.

Method used

By combining the attention mechanism and the A2C algorithm, a Markov decision process (MDP) model of radar, target and jammer is constructed. The state features are extracted by the self-attention mechanism, and reinforcement learning is carried out through the actor-critic architecture to optimize the frequency agility strategy and achieve high-precision and low-latency anti-jamming decision.

Benefits of technology

It improves the decision-making accuracy and learning efficiency of frequency agility strategies, enhances the anti-interference capability of radar systems in complex electromagnetic environments, enables adaptive adjustment of strategies to cope with dynamic interference, and improves the performance of radar systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669211B_ABST
    Figure CN120669211B_ABST
Patent Text Reader

Abstract

The application particularly relates to a radar intelligent anti-jamming method combining an attention mechanism and an A2C algorithm, which comprises the following steps: constructing a Markov decision process (MDP) model corresponding to a radar, a target and a jammer, and defining a state space, an action space and a multi-target reward function corresponding to the MDP model; performing feature extraction on state data by using a preprocessing layer based on a self-attention mechanism to obtain corresponding state features; wherein the state data comprises jamming signal frequency information; determining a frequency agility strategy corresponding to the state features by using an actor network; controlling the radar to execute the frequency agility strategy to collect updated state data; inputting the updated state data and the frequency agility strategy into a critic network to obtain a time difference deviation of an action value function under a current state; performing gradient updating on the actor network by using the time difference deviation, updating network parameters of the actor network by using the updated gradient, and generating the frequency agility strategy by using the actor network after the network parameters are updated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of radar technology, and specifically to a radar intelligent anti-jamming method that combines attention mechanism and A2C algorithm. Background Technology

[0002] Among related technologies, the development of radar anti-jamming technology stems from its application in the military field. In modern electronic warfare, radar undertakes crucial tasks such as detection and tracking, making its normal operation vital to the user. Therefore, Electronic Countermeasures (ECM) methods are widely used. Jammers can disrupt the enemy radar's effective use of the electromagnetic spectrum, thereby achieving interference. Under jamming conditions, radar cannot effectively receive information-carrying signals and may even be misled. To reduce the impact of jamming on radar, developing more convenient and effective radar anti-jamming technologies is particularly important in electronic warfare.

[0003] In electronic warfare, jamming techniques can be categorized into active and passive jamming based on whether the jamming equipment can actively emit electromagnetic waves. Active jamming, which involves using a jamming source to emit interference signals to disrupt the normal operation of radar, has received considerable attention. Currently, active jamming techniques have evolved into a variety of complex methods, such as frequency targeting jamming and frequency sweeping jamming. In practice, radars face complex environments with various types of interference, placing higher demands on the adaptability of radar anti-jamming technologies. Furthermore, the intelligence level of jammers is increasing. In some radar anti-jamming problems, jammers can utilize advanced methods to perform detailed analysis of radar anti-jamming strategies and improve their own interference signal transmission strategies. Therefore, addressing increasingly complex and intelligent jamming methods is a crucial research topic in radar anti-jamming.

[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] This invention provides a radar intelligent anti-jamming method combining attention mechanism and A2C algorithm, a computer program product, an electronic device, and a storage medium, which can adaptively adjust radar anti-jamming strategy and overcome the defects of the prior art to a certain extent.

[0006] Other features and advantages of the invention will become apparent from the following detailed description, or may be learned in part by practice of the invention.

[0007] According to a first aspect of the present invention, a radar intelligent anti-jamming method combining an attention mechanism and an A2C algorithm is provided, the method comprising:

[0008] Construct Markov Decision Process (MDP) models corresponding to radar, target, and jammer, and define the state space, action space, and multi-objective reward function corresponding to the MDP models;

[0009] A preprocessing layer based on a self-attention mechanism is used to extract features from the state data to obtain the corresponding state features; wherein, the state data includes interference signal frequency information;

[0010] The frequency agility strategy corresponding to the state characteristics is determined by using an actor network; and the radar is controlled to execute the frequency agility strategy in order to collect updated state data.

[0011] The updated state data and frequency agility strategy are input into the commentator network to obtain the time difference deviation of the action value function in the current state.

[0012] The actor network is updated with gradients using time difference bias, and the updated gradients are used to update the network parameters of the actor network, so as to generate a frequency agility strategy using the actor network with updated network parameters.

[0013] In some exemplary embodiments, the state space includes frequency information of radar transmitted signals and frequency information of interference signals, configured as observation information; the action space includes the radar frequency; and the multi-target reward function includes the radar detection probability and the radar signal-to-interference-plus-noise ratio.

[0014] In some exemplary embodiments, constructing the Markov Decision Process (MDP) model corresponding to the radar, target, and jammer includes:

[0015] Define the action space, including: at time t, the radar agent observes the current environment and obtains the corresponding state s. t And for state s t Determine the corresponding frequency agility strategy; the frequency agility strategy includes: carrier frequency a t The carrier frequency is selected from a given set of M frequency points.

[0016] Define the state space, including: the state s obtained by the radar agent from observing the current environment. t This includes frequency information of radar transmitted signals and frequency information of jamming signals transmitted by jammers;

[0017] Define a multi-objective reward function, including: the state s at time t. t Transition to the state s of the next moment t+1 The radar agent's received reward r t This is used to describe the effect of a radar agent's frequency agility strategy on radar anti-jamming.

[0018] In some exemplary embodiments, the frequency information of the interference signal includes a tuple based on the center frequency and bandwidth.

[0019] In some exemplary embodiments, the multi-objective reward function includes: a radar detection probability index and a radar signal-to-interference-plus-noise ratio index.

[0020] In some exemplary embodiments, a preprocessing layer based on a self-attention mechanism is used to extract features from the state data to obtain corresponding state features, including:

[0021] Map the state data into a query matrix Q, a key matrix K, and a value matrix V;

[0022] Configure the attention weight matrix based on the query matrix Q, the key matrix K, and the value matrix V;

[0023]

[0024] Where, d k Dimensions are matrices;

[0025] The state feature vector is calculated using the attention weight matrix; the state feature vector includes: based on the current frequency agile strategy π. θ In state s t Select radar frequency hopping action a t The radar agent performs frequency hopping actions. t It transmits signals, receives interference signals from jammers in the environment, and calculates the detection probability and signal-to-interference-plus-noise ratio based on the frequency information of radar and jammers in the environment, thus obtaining the return r. t and the state s at the next moment t+1 .

[0026] In some exemplary embodiments, updated state data and frequency agility strategies are input into the commentator network to obtain the time difference deviation of the action-value function in the current state, including:

[0027] δ t =r t +γV φ (s t+1 )-V φ (s t )

[0028] Among them, V φ (s t ) is the state value function;

[0029] Construct the MSE loss function and minimize it to update the network parameters.

[0030] In some exemplary embodiments, the method further includes: outputting the optimal action of the radar based on the current state of the radar and feedback from the commentator network, including:

[0031] Calculate the advantage function A φ (s t ,a t ) and use it to calculate the gradient of the actor network, including:

[0032]

[0033] Use gradients to update the actor network parameters, including:

[0034]

[0035] Among them, s t Let a be the state at time t. t Let t be the radar frequency modulation action at time t.

[0036] According to a second aspect of the present invention, a computer program product is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, the radar intelligent anti-jamming method combining the attention mechanism and the A2C algorithm described above is implemented.

[0037] According to a third aspect of the present invention, an electronic device is provided, comprising:

[0038] Processor; and

[0039] Memory for storing the executable instructions of the processor;

[0040] The processor is configured to implement the radar intelligent anti-jamming method combining the attention mechanism and the A2C algorithm described above when executing the executable instructions.

[0041] According to a fourth aspect of the present invention, a storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described radar intelligent anti-jamming method combining the attention mechanism and the A2C algorithm.

[0042] The radar intelligent anti-jamming method combining attention mechanism and A2C algorithm provided in the embodiments of the present invention can accurately extract key signals by setting a preprocessing layer based on self-attention mechanism, thereby effectively improving the decision accuracy of frequency agility strategy. The reinforcement learning method based on actor-critic architecture enables policy optimization and value estimation to be performed in parallel, resulting in faster convergence and avoiding the computational difficulties of traditional methods in high-dimensional state spaces. Simultaneously, the system continuously updates the strategy through a real-time feedback mechanism, enabling the radar to adapt to dynamically changing interference environments and improving anti-jamming capabilities. In summary, this solution has advantages over existing technologies in terms of intelligent decision-making for frequency agility, learning efficiency, and environmental adaptability, and can effectively improve the performance of radar systems in complex electromagnetic environments.

[0043] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0044] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0045] Figure 1 The diagram illustrates an exemplary embodiment of the present invention of a radar intelligent anti-jamming method combining an attention mechanism and an A2C algorithm;

[0046] Figure 2 The illustration shows a schematic diagram illustrating the principle of a radar intelligent anti-jamming method combining an attention mechanism and an A2C algorithm, as an exemplary embodiment of the present invention.

[0047] Figure 3 The illustration shows a schematic diagram of a radar intelligent anti-jamming method combining an attention mechanism and an A2C algorithm, according to an exemplary embodiment of the present invention.

[0048] Figure 4 This illustration schematically shows the performance of different algorithms when the reward is the signal-to-interference-plus-noise ratio, according to an exemplary embodiment of the present invention.

[0049] Figure 5 This illustration schematically shows the performance of different algorithms when the reward is the detection probability, according to an exemplary embodiment of the present invention.

[0050] Figure 6 This schematic diagram illustrates a convergence comparison when SINR is a reward, as an exemplary embodiment of the present invention.

[0051] Figure 7This diagram illustrates a convergence comparison of an exemplary embodiment of the present invention when the detection probability is a reward.

[0052] Figure 8 This schematic diagram illustrates the composition of an electronic device according to an exemplary embodiment of the present invention. Detailed Implementation

[0053] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the invention will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0054] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0055] In related technologies, traditional radar anti-jamming methods rely on fixed frequency hopping rules or simple reinforcement learning algorithms, which are insufficient to cope with the dynamic game-playing of intelligent jammers. Existing technologies suffer from high variance in policy gradient algorithms, while value function-based methods (such as DQN) are inefficient in continuous action spaces. Furthermore, radar signal feature extraction lacks the ability to focus on key jamming signals, resulting in insufficient decision robustness.

[0056] To address the shortcomings and deficiencies of existing technologies, this example implementation provides a radar intelligent anti-jamming method combining an attention mechanism and an A2C algorithm. By dynamically modeling the game process between the radar and the jammer, enhancing the ability to identify key signals, and optimizing the strategy update mechanism, it achieves high-precision, low-latency anti-jamming decision-making. Specifically, refer to... Figure 1 As shown, the radar intelligent anti-jamming method combining attention mechanism and A2C algorithm can specifically include the following steps:

[0057] Step S11: Construct Markov Decision Process (MDP) models corresponding to the radar, target, and jammer, and define the state space, action space, and multi-objective reward function corresponding to the MDP models.

[0058] Step S12: Use a preprocessing layer based on a self-attention mechanism to extract features from the state data and obtain the corresponding state features; wherein, the state data includes interference signal frequency information;

[0059] Step S13: Use the actor network to determine the frequency agility strategy corresponding to the state characteristics; and control the radar to execute the frequency agility strategy to collect updated state data.

[0060] Step S14: Input the updated state data and frequency agility strategy into the commentator network to obtain the time difference deviation of the action value function in the current state;

[0061] Step S15: Update the gradient of the actor network using the time difference bias, and update the network parameters of the actor network using the updated gradient, so as to generate a frequency agility strategy using the actor network with updated network parameters.

[0062] The following will describe in more detail each step of the radar intelligent anti-jamming method combining attention mechanism and A2C algorithm in this exemplary embodiment, with reference to the accompanying drawings and embodiments.

[0063] In step S11, a Markov Decision Process (MDP) model corresponding to the radar, target, and jammer is constructed, and the state space, action space, and multi-objective reward function corresponding to the MDP model are defined.

[0064] For example, the state space includes frequency information of radar transmitted signals and frequency information of interference signals, configured as observation information; the action space includes the radar frequency; the multi-target reward function includes: the radar detection probability and the radar signal-to-interference-plus-noise ratio.

[0065] For example, constructing the Markov Decision Process (MDP) model corresponding to the radar, target, and jammer includes:

[0066] Define the action space, including: at time t, the radar agent observes the current environment and obtains the corresponding state s. t And for state s t Determine the corresponding frequency agility strategy; the frequency agility strategy includes: carrier frequency a t The carrier frequency is selected from a given set of M frequency points.

[0067] Define the state space, including: the state s obtained by the radar agent from observing the current environment. t This includes frequency information of radar transmitted signals and frequency information of jamming signals transmitted by jammers;

[0068] Define a multi-objective reward function, including: the state s at time t. t Transition to the state s of the next momentt+1 The radar agent's received reward r t This is used to describe the effect of a radar agent's frequency agility strategy on radar anti-jamming.

[0069] For example, the frequency information of the interference signal includes a tuple based on the center frequency and bandwidth.

[0070] Specifically, we can first define a Markov decision process model for radar anti-jamming. Specifically, at time t, the radar obtains its state s at time t by observing the environment. t The radar agent selects an action, that is, it selects the current state s based on the current policy π. t The corresponding carrier frequency a t It also transmits radar radio frequency signals. Simultaneously, the jammer can be considered part of the environment; it is capable of transmitting jamming signals. Interference signal It is a binary tuple including center frequency and bandwidth. Then the state will change from s t Transfer to the next moment s t+1 The radar agent will receive a reward r t It describes the advantages and disadvantages of the radar agent's choice of frequency hopping action for radar anti-jamming effect.

[0071] Based on the construction of the radar anti-jamming model, the key elements contained in the MDP can be defined in the reinforcement learning algorithm for radar anti-jamming as follows:

[0072] The action space of the model describes the radar's actions, specifically its frequency agility according to a strategy. For each radar pulse, its carrier frequency can be arbitrarily chosen from a given set of M frequency points. Here, M represents the size of the action space in reinforcement learning. Since the frequency of the frequency-agile radar's transmitted waveform changes linearly with time, the agent's action can be set to select a suitable frequency from the M frequency points. This can be represented as:

[0073]

[0074] in, It is the starting frequency of the LFM waveform (linear frequency modulation waveform) transmitted by the radar. This is the end frequency of the LFM waveform, measured in Hertz.

[0075] For the model's state space, consider the use of interference detectors and other methods to sense interference signals during radar operation. In this process, the radar's perception of the interference signals is configured as its observation result. By sensing and analyzing interference signals, the radar agent can gain a more accurate understanding of the current electromagnetic environment, providing crucial reference information for formulating and adjusting anti-interference strategies. The current state in the radar's anti-interference environment can be defined as:

[0076]

[0077] Where i∈{0,1,…,t}, and These are the frequency information of the signals transmitted by the radar and the jammer, respectively.

[0078] To enhance the information acquired by the radar agent and learn from historical experience, the agent can utilize historical information H. t It replaces the current state information to learn and make decisions. In RL (Reinforcement Learning) theory, historical information H... t It can be viewed as a result of action a t and observation o t A sequence that constitutes this is represented as:

[0079] H t =a0,o1,…,a t-1 ,o t

[0080] Among them, o t This refers to the radar's observations of interference signals, specifically the frequency information of those signals. In practical applications, if the radar makes decisions based solely on data from its interaction with the environment, then the historical data contains all the information about the environment. t This refers to the frequency information of the radar's transmitted signals.

[0081] Clearly, as time t increases, H t The size of H also increases significantly, which makes it difficult to input the model. Therefore, radar cannot directly use it as input state. To solve this problem, the kth-orderhistory method can be used to obtain an approximate result of historical information. This method uses a sequence of the past k observations and actions to approximate the historical information H. t Therefore, the state of the frequency-agile radar agent at time t can be represented as:

[0082] s t =a t ,o t ,a t-1 ,…,ot-k

[0083] in, This represents the carrier information of the radar and jammer at time t, that is, the start and end frequencies of one cycle, representing the transmitted signal information. t This refers to the radar's action at time t, specifically the frequency information of the radar's transmitted signal. The radar has m selectable frequency bands. Since linear frequency modulated signals have a start and end frequency, the action space is 1×m×m. The radar's observation space has a dimension of 3×k, where the first dimension is 3. k is the length of the extracted historical information.

[0084] The rewards for the model can be used to evaluate the effectiveness of the agent's strategy. The purpose of radar anti-jamming is to ensure the normal operation of the radar and to complete tasks such as detection. Therefore, the formulation of rewards should also follow the principle of reflecting the radar's anti-jamming capability. This chapter establishes evaluation criteria for the performance of radar anti-jamming strategies, mainly including two indicators: radar detection probability and radar signal-to-interference-plus-noise ratio (SNR).

[0085] The detection probability of a radar system refers to the probability that the system can correctly detect a target in the presence of noise interference and a target object. A high detection probability means that the radar can still effectively detect targets under various interference environments, reflecting the reliability and sensitivity of the radar system. Generally speaking, when the false alarm probability of a frequency-agile radar is set to a very small value and a single-pulse detection method is used, the detection probability P of the frequency-agile radar is... d It can be calculated using the following formula, which is expressed as:

[0086]

[0087] Among them, S N It is the signal-to-noise ratio of a single pulse, y0 = -ln(P f ), P f It is the probability of a false alarm.

[0088] When the false alarm probability P of frequency-agile radar f =10 -6 Furthermore, when the pulse accumulation number n received by the frequency-agile radar is very large, the detection probability of the target by the frequency-agile radar can be calculated using the following formula:

[0089]

[0090] Where n is the accumulated radar pulse value in one scan, s N This is the signal-to-noise ratio result of a single pulse of the radar. Among them,

[0091]

[0092] In the formula for calculating the detection probability, D M =n γ γ is the accumulation efficiency, and its value is related to the size of n.

[0093] In most cases, the value of γ is between 0.7 and 0.9, and only when it is very large will its value approach 0.5. f This is the upper α quantile of the false alarm probability.

[0094] The signal-to-interference-plus-noise ratio (SIR / NNR) of a radar is the ratio of the radar signal to the sum of interference and noise. In practical terms, it refers to the ratio between the strength of the useful signal received by the radar and the strength of the interference signal (noise plus interference) received by the radar. The formula is expressed as:

[0095]

[0096] Signal-to-interference-plus-noise ratio (SINNR) is an important indicator for measuring the extent of interference to a radar. The higher the SINNR, the less interference the radar experiences and the better its anti-jamming performance.

[0097] These two metrics allow us to comprehensively evaluate the performance of radar anti-jamming strategies. During the learning process, the agent continuously optimizes the frequency hopping strategy to maximize the detection probability and signal-to-interference-plus-noise ratio, thereby improving the radar system's anti-jamming capability in complex electromagnetic environments.

[0098] In step S12, a preprocessing layer based on a self-attention mechanism is used to extract features from the state data to obtain the corresponding state features; wherein, the state data includes interference signal frequency information.

[0099] For example, state data is mapped to a query matrix Q, a key matrix K, and a value matrix V;

[0100] Configure the attention weight matrix based on the query matrix Q, the key matrix K, and the value matrix V;

[0101]

[0102] Where, d k Dimensions are matrices;

[0103] The state feature vector is calculated using the attention weight matrix; the state feature vector includes: based on the current frequency agile strategy π. θ In state s t Select radar frequency hopping action a t The radar agent performs frequency hopping actions. tIt transmits signals, receives interference signals from jammers in the environment, and calculates the detection probability and signal-to-interference-plus-noise ratio based on the frequency information of radar and jammers in the environment, thus obtaining the return r. t and the state s at the next moment t+1 .

[0104] Specifically, the calculation formula for the self-attention mechanism includes:

[0105]

[0106] Where Q = XW Q K = XW K V = XW V X is the input feature, W Q W K and W V It is a trainable parameter matrix.

[0107] Since the input feature form is s t =a t ,o t-1 ,a t-1 ,…,o t-k ], k = 4. Therefore, we can obtain:

[0108] Q = s t W Q

[0109] K = s t W K

[0110] V = s t W V

[0111] The preprocessing layer based on the self-attention mechanism outputs a weighted state feature vector, including the features based on the current frequency hopping strategy π. θ In state s t Select radar frequency hopping action a t The radar agent performs frequency hopping actions. t The system transmits signals, jammers in the environment transmit jamming signals, and the detection probability and signal-to-interference-plus-noise ratio are calculated based on the frequency information of the radar and jammers in the environment to obtain the return r. t and the state s at the next moment t+1 .

[0112] Specifically, by introducing a self-attention-based preprocessing layer before the first layer of the actor network in the Attention-A2C algorithm, the state (including radar and environmental information) is processed. This self-attention layer allows the model to better focus on important parts of the state space. This not only improves the model's performance and generalization ability but also helps the model capture long-range dependencies in the state space. Specifically, the self-attention layer can better understand and process sequential data by processing state changes in historical radar information.

[0113] In step S13, the frequency agility strategy corresponding to the state characteristics is determined using the actor network; and the radar is controlled to execute the frequency agility strategy to collect updated state data.

[0114] For example, in the actor network, the optimal action of the radar can be output based on the current state of the radar and the feedback from the commentator network, i.e., the frequency agility strategy.

[0115] Specifically, the radar agent observes and acquires jammer frequency information, combines this information with its own actions to form an action space, represented as:

[0116]

[0117] Constructing radar frequency hopping action And observation of jammers Select the optimal frequency agility strategy from a predefined set of candidate frequency points and perform frequency hopping.

[0118] In step S14, the updated state data and frequency agility strategy are input into the commentator network to obtain the time difference deviation (TD) of the action value function in the current state.

[0119] For example, for a commentator network, a value function can be constructed by estimating the radar's long-term cumulative reward based on the radar's current state and actions.

[0120] Specifically, in a critic network, the calculation of the Q-value can be divided into two parts: the state-value function V(s) and the dominance value A(s,a). The formula is expressed as:

[0121] Q(s,a)=V(s)+A(s,a)

[0122] The A2C algorithm uses the advantage function A. π (s t ,a t This function, called , is used to replace the value function in the critic network. Its advantage lies in its ability to be used to evaluate the quality of the selected agent's action value by comparing it to the average of all candidate actions.

[0123] The advantage function can be defined as:

[0124]

[0125] Among them, Q π (s t ,a t ) is the action-value function, V π (s t V is the state-value function. Using an advantage function instead of a value function avoids penalizing certain actions due to the agent's current poor policy. Similarly, it avoids situations where an advantage doesn't reward actions when the agent is currently following a good policy. Where V... π (s t The unknown value (V) requires a parameterized network to estimate it, which can be denoted as V. φ (s), at this point the formula for the policy gradient in A2C becomes:

[0126]

[0127] Here, A φ (s t ,a t The state value function V can be used. φ (s t Estimate and compute. Similar to the AC algorithm, V in A2C... φ (s t It is also updated using the time-difference method, and the TD objective here is:

[0128] y = r t +γV φ (s t+1 )

[0129] The TD deviation is expressed as follows:

[0130] δ t =r t +γV φ (s t+1 )-V φ (s t )

[0131] The parameters of a value network can be updated by constructing and minimizing the MSE loss function.

[0132] A φ (s t ,a tWhen the value of θ is greater than 0, it indicates that the selected action is better than the average action. In this case, its value is positive, which allows the network parameter θ to move along the positive gradient direction. Conversely, when its value is less than 0, it means that the agent's action is worse than the average action, and this situation should be avoided. A negative value allows the network parameter θ to move along the negative gradient direction.

[0133] In step S15, the actor network is updated with gradients using time difference bias, and the updated gradients are used to update the network parameters of the actor network, so as to generate a frequency agility strategy using the actor network with updated network parameters.

[0134] For example, the actor network can output the optimal action for the radar based on the radar's current state and feedback from the commentator network. Specifically, the advantage function A can be calculated. φ (s t ,a t And use it to calculate the gradient of the actor network:

[0135]

[0136] Then update the actor network parameters using gradients: This is represented as:

[0137]

[0138] Specifically, the radar executes corresponding frequency agility actions based on the output of the actor network, while simultaneously observing environmental feedback and updating the radar's status and rewards. This process is repeated cyclically to achieve intelligent anti-interference for the radar.

[0139] For example, to demonstrate the effectiveness of the present invention, the following simulation experiments are used for further illustration.

[0140] The dataset was used for radar modeling and simulation based on the radarsimpy library. We performed detailed modeling of both the radar and the jammer, allowing adjustment of parameters such as power and frequency to achieve interaction under different conditions. In the simulation verification, the radar's center frequency was set to 6 GHz, varying in steps of Δf = 50 MHz. The radar could perform frequency agility across 4–12 sub-bands, selecting an appropriate frequency band to output the radar signal. The radar transmitted N = 4 pulses in a pulse sequence, with each pulse lasting 8 × 10⁻⁶ pulses. -6 The radar power is set to 20dBm, and the jammer power is set to 30dBm. The radar receiver noise figure is 2dB, and the total baseband gain is 60dB.

[0141] The technical solution of this invention is compared with various existing technical standards in radar anti-jamming models, including the Advantageous Actor-Critic Algorithm (A2C), Deep Q Network (DQN), Dual Deep Q Network (DDQN), Policy Gradient Learning (SARSA), and Random Policy.

[0142] Prediction metrics are defined. This invention utilizes commonly used metrics in its radar anti-jamming strategy to evaluate the performance of all models, including detection probability and signal-to-interference-plus-noise ratio (SINR). These metrics are visually reflected through different action spaces. Furthermore, the convergence of this invention and A2C is compared. In the continuous action space, the size of the action space represents the range of selectable frequency points for the radar agent, and the jamming method is frequency-based interference. Detection probability refers to the probability that the radar correctly detects the target, i.e., the likelihood of the radar successfully identifying the target given its presence. It is typically affected by factors such as signal strength, noise level, and interference intensity; a higher detection probability indicates stronger radar detection capability. SINR represents the ratio of the received signal strength to the interference and noise.

[0143] refer to Figures 4-7 The figure shows the preliminary verification comparison results between the present invention and existing advanced technologies in this research field.

[0144] refer to Figure 4 As shown, this diagram illustrates the performance of different algorithms under frequency interference conditions, with signal-to-interference-plus-noise ratio (SNR) and detection probability as rewards. The horizontal axis represents the action space of the radar agent, and the vertical axis represents SNR and detection probability, respectively. As the action space increases, the radar's anti-jamming performance improves. When the reward is SNR, the performance difference between the Attention-A2C algorithm and several reinforcement learning algorithms such as DQN and DDQN is not significant when the action space is small. However, as the action space increases, the Attention-A2C algorithm and the A2C algorithm significantly outperform other algorithms, indicating that these algorithms perform better in a larger action space.

[0145] When the reward is the detection probability, the Attention-A2C and A2C algorithms outperform other algorithms even in a small action space, and maintain their advantage as the action space increases. The results of DQN and DDQN are similar, slightly higher than the SARSA algorithm. This indicates that although DQN and DDQN handle interference well, the Attention-A2C and A2C algorithms still have a significant advantage in improving the detection probability.

[0146] The Attention-A2C algorithm performs excellently in aiming interference environments with varying action spaces, demonstrating strong adaptability. Its introduced attention mechanism not only improves the algorithm's performance but also enhances its ability to perform in large action spaces.

[0147] Besides the algorithm's reward, convergence is also a crucial indicator of its effectiveness. Convergence primarily measures the speed and stability with which an algorithm approaches the optimal solution during iteration. Better convergence indicates higher algorithm usability, as it signifies that the algorithm can find the optimal solution faster and more stably.

[0148] refer to Figures 5-7 As shown, the optimization effect of the attention mechanism on the convergence of this invention and the A2C algorithm is compared and analyzed. The performance of the A2C and Attention-A2C algorithms is compared over 300 training rounds. The figure shows the changes in signal-to-noise ratio and detection probability with training rounds.

[0149] In terms of convergence speed, when using detection probability as the reward, the Attention-A2C algorithm converged after 60 training episodes, significantly faster than the A2C algorithm which required 100 training episodes. The Attention-A2C algorithm also converged faster than the A2C algorithm in terms of both signal-to-noise ratio and detection probability.

[0150] In terms of convergence stability, both algorithms exhibit good stability. However, in experiments where signal-to-noise ratio is used as a reward, the Attention-A2C algorithm performs better. Its introduced attention mechanism enables the radar agent to more accurately focus on key states and signals, thereby improving learning efficiency and stability.

[0151] In summary, the addition of the attention mechanism not only significantly improves the convergence speed of the radar anti-jamming algorithm, but also enhances its convergence stability.

[0152] The method provided in the embodiments of the present invention is referred to Figure 3As shown, compared to existing technologies, this solution possesses stronger adaptability and optimization efficiency. Traditional methods often rely on fixed rules or heuristic algorithms, making it difficult to flexibly cope with complex interference environments. This invention, however, constructs an MDP model through reinforcement learning, enabling the radar to dynamically perceive the environment and adaptively adjust its frequency-agile strategy. Furthermore, existing methods often employ simple weighted averaging when processing historical radar observation information, resulting in low information utilization. This invention introduces a self-attention mechanism to accurately extract key signals, improving decision-making accuracy. The reinforcement learning method based on the actor-critic architecture allows policy optimization and value estimation to proceed in parallel, leading to faster convergence and avoiding the computational difficulties of high-dimensional state spaces inherent in traditional methods. Simultaneously, the system continuously updates its strategy through a real-time feedback mechanism, enabling the radar to adapt to dynamically changing interference environments and enhancing its anti-jamming capabilities. In summary, this solution demonstrates superior advantages over existing technologies in terms of intelligent decision-making with frequency agility, learning efficiency, and environmental adaptability, effectively improving the performance of radar systems in complex electromagnetic environments.

[0153] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may, for example, be executed synchronously or asynchronously in multiple modules.

[0154] It should be noted that although several modules or units of the device for performing actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0155] Figure 8 A schematic diagram of an electronic device suitable for implementing embodiments of the present invention is shown.

[0156] It should be noted that, Figure 8 The electronic device 1000 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0157] like Figure 8As shown, the electronic device 1000 includes a Central Processing Unit (CPU) 1001, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1002 or programs loaded from Storage Unit 1008 into Random Access Memory (RAM) 1003. The RAM 1003 also stores various programs and data required for system operation. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An Input / Output (I / O) interface 1005 is also connected to the bus 1004.

[0158] The following components are connected to I / O interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to I / O interface 1005 as needed. Removable media 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1010 as needed so that computer programs read from them can be installed into storage section 1008 as needed.

[0159] In particular, according to embodiments of the present invention, the processes described below with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a storage medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1009, and / or installed from removable medium 1011. When the computer program is executed by central processing unit (CPU) 1001, it performs various functions defined in the system of this application.

[0160] It should be noted that the storage medium shown in the embodiments of the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, wherein computer-readable program code is carried. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0162] The units described in the embodiments of the present invention can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0163] It should be noted that, as another aspect, this application also provides a storage medium, which may be included in an electronic device or may exist independently without being assembled into the electronic device. The aforementioned storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to perform the methods described in the following embodiments. For example, the electronic device may perform... Figure 1 The steps of the method shown.

[0164] In one embodiment, this application provides a computer program product including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0165] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0166] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims.

[0167] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A radar intelligent anti-jamming method combining attention mechanism and A2C algorithm, characterized in that, The method includes: A Markov Decision Process (MDP) model corresponding to the radar, target, and jammer is constructed, and the state space, action space, and multi-target reward function corresponding to the MDP model are defined. The state space includes the frequency information of the radar transmitted signal and the frequency information of the jamming signal, which are configured as observation information. The action space includes the radar frequency. The multi-target reward function includes the radar detection probability and the radar signal-to-interference-plus-noise ratio. The construction of the Markov Decision Process (MDP) model corresponding to the radar, target, and jammer includes: Define the action space, including: at time t, the radar agent observes the current environment and obtains the corresponding state. and in response to the status Determine the corresponding frequency agility strategy; the frequency agility strategy includes: carrier frequency The carrier frequency is from a given Selected from a number of frequency points; Define the state space, including: the state obtained by the radar agent from observing the current environment. This includes frequency information of radar transmitted signals and frequency information of jamming signals transmitted by jammers; Define a multi-objective reward function, including: the state at time t. Transition to the state of the next moment The radar agent receives rewards , used to describe the effect of radar agents selecting frequency agility strategies on radar anti-jamming; A preprocessing layer based on a self-attention mechanism is used to extract features from the state data to obtain the corresponding state features, including: mapping the state data to a query matrix Q, a key matrix K, and a value matrix V; and configuring an attention weight matrix based on the query matrix Q, the key matrix K, and the value matrix V. in, Dimensions are matrices; The state feature vector is calculated using the attention weight matrix after weighting the state data; the state feature vector includes: based on the current frequency agile strategy. In state Select radar frequency hopping action The radar agent performs frequency hopping. It transmits signals, receives interference signals emitted by jammers in the environment, and calculates the detection probability and signal-to-interference-plus-noise ratio based on the frequency information of radar and jammers in the environment, thus obtaining a return. and the state at the next moment The status data includes interference signal frequency information. The frequency agility strategy corresponding to the state characteristics is determined by using an actor network; and the radar is controlled to execute the frequency agility strategy in order to collect updated state data. The updated state data and frequency agility strategy are input into the commentator network to obtain the time difference deviation of the action value function in the current state. The actor network is updated with gradients using time difference bias, and the updated gradients are used to update the network parameters of the actor network, so as to generate a frequency agility strategy using the actor network with updated network parameters.

2. The method according to claim 1, characterized in that, The frequency information of the interference signal includes a binary tuple based on the center frequency and bandwidth.

3. The method according to claim 1, characterized in that, The multi-objective reward function includes: the radar detection probability index and the radar signal-to-interference-plus-noise ratio index.

4. The method according to claim 1, characterized in that, The updated state data and frequency agility strategy are input into the commentator network to obtain the time difference bias of the action-value function in the current state, including: in, It is a state-value function; Construct the MSE loss function and minimize it to update the network parameters.

5. The method according to claim 1, characterized in that, The method further includes: outputting the optimal action of the radar based on the current state of the radar and feedback from the commentator network, including: Calculate the advantage function And use it to calculate the gradient of the actor network; Use gradients to update the actor network parameters, including: in, The state at time t, Let t be the radar frequency modulation action at time t.

6. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the radar intelligent anti-jamming method combining the attention mechanism and the A2C algorithm as described in any one of claims 1 to 5.

7. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the radar intelligent anti-jamming method combining the attention mechanism and the A2C algorithm as described in any one of claims 1 to 5 by executing the executable instructions.