Multifunctional radar interference strategy generation method and device, equipment and medium
By combining partially visible Markov decision model and Transformer-A2C network, an interference strategy generation network is built, which solves the problem of insufficient exploration of response delay and parameter combination of existing radar jamming strategy generation methods when facing multifunctional radars, and real-time and effective interference strategy generation is achieved.
Patent Information
- Application Number
- CN202510427767.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-22
AI Technical Summary
The existing radar jamming strategy generation method is not effective when facing anti-jamming technologies such as pulsed Doppler of modern multifunctional radars, and has response delays, making it difficult to cope with dynamic changes in the radar working mode, and it is impossible to effectively explore all parameter combinations.
A partially considerable Markov decision model and interference strategy combined with Transformer and A2C network are used to generate a network. The TIT feature extraction layer, the policy network generates an action probability distribution layer and the value network evaluates the state value layer, forming a dual-branch decoupling structure to generate interference strategies in real time.
It realizes that in a complex and dynamic radar countermeasure environment, it can quickly respond and generate effective jamming strategies, which improves the intelligence and flexibility of radar interference decisions and adapts to the rapid changes in radar state.
Smart Images

Figure CN120352839A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of radar electronic countermeasure technology, and particularly to a method, device, equipment and medium for generating a multi-functional radar jamming strategy. Background Art
[0002] Traditional electronic countermeasure jamming strategies usually rely on a large amount of prior knowledge and expert experience, comparing radar parameters with a jamming knowledge base through template matching to select the best jamming strategy. However, with the progress of radar technology, modern multi-functional radars have adopted anti-jamming technologies such as pulse Doppler, intra-pulse parameter agility, and pulse compression, significantly enhancing their anti-jamming capabilities. At this time, static jamming strategies are becoming increasingly ineffective and may even be recognized by the radar, resulting in poor jamming effects. Therefore, optimizing jamming decisions has become particularly important, which requires the jamming strategy generation method to not only select appropriate jamming modes according to the radar state but also respond quickly when the radar state changes to maintain continuous protection of the target. In this context, sequential decision-making methods emphasize how the jammer reasonably allocates jamming resources after perceiving the electromagnetic environment and continuously learns and updates strategies to maximize the jamming effect and cope with changes in the radar state. As a dynamic decision-making framework, reinforcement learning provides new ideas for radar jamming decisions. Through the interaction between the agent and the environment, reinforcement learning can continuously optimize strategies, adapt to complex and dynamic confrontation environments, and demonstrate relatively excellent adaptive capabilities.
[0003] However, existing reinforcement learning methods such as Q-learning, SARSA, and DQN have some limitations in radar jamming decisions. First, Q-learning and SARSA are usually optimized for a single radar operating mode or fixed parameters, which makes them perform poorly in dealing with the dynamic changes of radar switching operating modes. Second, although DQN can handle more complex state spaces, it may be limited by discretization in continuous parameter spaces, resulting in an inability to effectively explore all possible parameter combinations. Finally, these algorithms often have response delays in the real-time decision-making process, especially in the case of rapid changes in the radar operating state, and may not be able to make optimal decisions in a timely manner. Summary of the Invention
[0004] Based on this, it is necessary to provide a multi-functional radar jamming strategy generation method, device, equipment and medium that can timely generate radar jamming strategies for the current situation in response to the above technical problems.
[0005] A multi-functional radar jamming strategy generation method, the method includes:
[0006] Obtain a training data set, where the training data set includes the characteristics of various radar signals and corresponding operating state labels;
[0007] Using a partially observable Markov decision model, the non - cooperative game process of multi - function radar countermeasure is modeled according to the "OODA" loop to obtain a multi - function radar countermeasure process model, and a jamming strategy generation network is constructed by integrating Transformer and A2C network. The jamming strategy generation network includes a TIT feature extraction layer, a policy network for generating action probability distribution, and a value network for evaluating state value, thus forming a double - branch decoupled structure;
[0008] Under the multi - function radar countermeasure process model, the jamming strategy generation network is trained using the training data set to obtain a trained jamming strategy generation network;
[0009] The real - time radar signal emitted by the jamming target is obtained, and the features of the real - time radar signal are input into the trained jamming strategy generation network to generate a corresponding jamming strategy.
[0010] In one embodiment, each of the TIT units includes an internal transformer and an external transformer. In the l - th TIT unit:
[0011]
[0012] In the above formula, z l-1 represents the output data of the internal transformer of the previous TIT unit, represents the internal transformer, represents the concatenation of all z l [0] across K time steps, and L represents the number of TIT units.
[0013] In one embodiment, both the policy network and the value network are constructed by vertically stacking multiple TIT units, a feed - forward neural network, and a Softmax layer.
[0014] In one embodiment, the multi - function radar countermeasure process model includes: the working state space of the multi - function radar, the jamming action space, the state transition probability, the reward function, the observation space, and the discount factor.
[0015] In one embodiment, the working state space includes: coarse search working state, precise search working state, detection working state, intercept working state, range resolution working state, stable tracking working state, and passive tracking working state;
[0016] The jamming action space includes: dense false target jamming, Doppler scintillation jamming, noise repetition jamming, chaff jamming, velocity - distance gate pull - off, and cross - eye jamming.
[0017] In one embodiment, the reward function is expressed as:
[0018]
[0019] r t = R base (s t+1 |s t ,a t ) + R add (s t+1 |s t ,a t )
[0020] In the above formula, T(s t ) represents the threat level when the multi-functional radar is in state s t , and the excitation value r t consists of a basic excitation R base and an additional excitation R add in two parts.
[0021] In one embodiment, the characteristics of the radar signal are that a radar phrase is formed by the working parameters of the multi-functional radar seeker.
[0022] This application also provides a device for generating a multi-functional radar jamming strategy, and the device includes:
[0023] A training data integration module, configured to obtain external data to obtain a training data set, where the training data set includes the characteristics of various radar signals and corresponding working state labels;
[0024] A non-cooperative game process modeling and interference strategy generation network construction module, configured to use a partially observable Markov decision model to model the non-cooperative game process of multi-functional radar countermeasure according to the "OODA" loop to obtain a multi-functional radar countermeasure process model, and fuse Transformer and A2C network to construct an interference strategy generation network, and the interference strategy generation network includes a TIT feature extraction layer, a policy network to generate an action probability distribution layer, and a value network to evaluate the state value layer, so as to form a double-branch decoupled structure;
[0025] An interference strategy generation network training module, configured to train the interference strategy generation network by using the training data set under the multi-functional radar countermeasure process model to obtain a trained interference strategy generation network;
[0026] An interference strategy real-time generation module, configured to obtain the real-time radar signal emitted by the interference target and input the characteristics of the real-time radar signal into the trained interference strategy generation network to generate a corresponding interference strategy.
[0027] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented:
[0028] Obtain a training data set, which includes the characteristics of various radar signals and the corresponding working state labels;
[0029] Adopt a partially observable Markov decision model to model the non - cooperative game process of multi - functional radar countermeasure according to the "OODA" loop, obtain a multi - functional radar countermeasure process model, and fuse Transformer and A2C network to construct an interference strategy generation network. The interference strategy generation network includes a TIT feature extraction layer, a policy network for generating action probability distributions, and a value network for evaluating state values, thus forming a double - branch decoupled structure;
[0030] Under the multi - functional radar countermeasure process model, use the training data set to train the interference strategy generation network to obtain a trained interference strategy generation network;
[0031] Obtain the real - time radar signal emitted by the interference target, and input the characteristics of the real - time radar signal into the trained interference strategy generation network to generate the corresponding interference strategy.
[0032] A computer - readable storage medium stores a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0033] Obtain a training data set, which includes the characteristics of various radar signals and the corresponding working state labels;
[0034] Adopt a partially observable Markov decision model to model the non - cooperative game process of multi - functional radar countermeasure according to the "OODA" loop, obtain a multi - functional radar countermeasure process model, and fuse Transformer and A2C network to construct an interference strategy generation network. The interference strategy generation network includes a TIT feature extraction layer, a policy network for generating action probability distributions, and a value network for evaluating state values, thus forming a double - branch decoupled structure;
[0035] Under the multi - functional radar countermeasure process model, use the training data set to train the interference strategy generation network to obtain a trained interference strategy generation network;
[0036] Obtain the real - time radar signal emitted by the interference target, and input the characteristics of the real - time radar signal into the trained interference strategy generation network to generate the corresponding interference strategy.
[0037] The above-mentioned multi-functional radar interference strategy generation method, device, computer equipment and storage medium model the non-cooperative game process of multi-functional radar countermeasure according to the "OODA" loop by adopting a partially observable Markov decision model, obtaining a multi-functional radar countermeasure process model, and designing an interference strategy generation network integrating Transformer and A2C, which includes a TIT feature extraction layer, a policy network for generating action probability distributions, and a value network for evaluating state values, thus forming a double-branch decoupled structure. Under the multi-functional radar countermeasure process model, the interference strategy generation network is trained using a training data set to obtain a trained interference strategy generation network. The features of the real-time radar signal emitted by the interference target are input into the trained interference strategy generation network to generate corresponding interference strategies. Using this method, effective and accurate interference strategies can be generated in real time. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 FIG. is a schematic flowchart of a multi-functional radar interference strategy generation method in an embodiment;
[0039] Figure 2 FIG. is a schematic diagram of possible transition situations of the working state of a multi-functional radar under ECM in an embodiment;
[0040] Figure 3 FIG. is a schematic diagram of convergence after training using the A2C reinforcement learning algorithm in a simulation experiment;
[0041] Figure 4 FIG. is a structural block diagram of a multi-functional radar interference strategy generation device in an embodiment;
[0042] Figure 5 FIG. is an internal structure diagram of computer equipment in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0044] In the existing technology, traditional electronic countermeasure interference strategies rely on prior knowledge and expert experience, and select strategies by template matching. Facing anti-jamming technologies such as pulse Doppler of modern multi-functional radars, static interference strategies have poor effects and may even be recognized by the radar. Using reinforcement learning can indeed optimize strategies to adapt to complex environments. However, most existing reinforcement learning methods have defects. For example, Q-learning and SARSA optimize for single or fixed parameters and are difficult to cope with the dynamic changes of radar working modes. DQN is restricted by the discretization of continuous parameter spaces and cannot effectively explore all parameter combinations. Moreover, these algorithms have response delays during real-time decision-making and cannot give optimal decisions in a timely manner when the radar working state changes rapidly. In this application, as Figure 1 shown, a method for generating a multi-functional radar interference strategy is provided, including the following steps:
[0045] Step S100, obtain a training data set, where the training data set includes the characteristics of various radar signals and corresponding working state labels.
[0046] Step S110, adopt a partially observable Markov decision model to model the non-cooperative game process of multi-functional radar countermeasure according to the "OODA" loop, obtain a multi-functional radar countermeasure process model, and fuse Transformer and A2C network to construct an interference strategy generation network. The interference strategy generation network includes a TIT feature extraction layer, a policy network for generating action probability distributions, and a value network for evaluating state values, thus forming a double-branch decoupled structure.
[0047] Step S120, under the multi-functional radar countermeasure process model, use the training data set to train the interference strategy generation network to obtain a trained interference strategy generation network.
[0048] Step S130, obtain the real-time radar signal emitted by the interference target, and input the characteristics of the real-time radar signal into the trained interference strategy generation network to generate a corresponding interference strategy.
[0049] In this application, the A2C reinforcement learning algorithm combines the advantages of policy gradients and value functions, and separately processes action selection and state evaluation through independent policy networks and value networks. Compared with Q-learning and SARSA, A2C can better handle continuous action spaces and can more efficiently cope with the state changes of radars in dynamic environments. Compared with DQN, A2C avoids complex technical problems through parallel training and a stable policy update mechanism, making the training process more efficient and converging faster. In addition, A2C has high stability and convergence during the decision-making process, can adapt to environmental changes faster, and thus provides a more intelligent and flexible solution for complex radar interference decisions.
[0050] Furthermore, existing methods often rely on static strategies or prior knowledge and lack the adaptive ability to dynamically change the radar state, resulting in unsatisfactory interference effects. By constructing a radar interference model based on the Partially Observable Markov Decision Process (POMDP) and combining it with the A2C algorithm, this method can automatically adjust the interference strategy in complex and dynamic combat environments to optimize the interference effect. Thus, more intelligent, flexible, and efficient electronic countermeasure interference decisions can be achieved, providing a new and effective solution for multifunctional radar countermeasures.
[0051] In this embodiment, considering the specific scenario of multifunctional radar countermeasures, that is, an electronic countermeasure aircraft flying with the target, in order to reduce the detection of our target by a multifunctional radar of the opposing party, the electronic countermeasure aircraft implements interference on this multifunctional radar. Therefore, in the following content, the multifunctional radar of the opposing party is regarded as the target, that is, the target of interference by the friendly electronic countermeasure aircraft.
[0052] Furthermore, in the specific scenario of the above-mentioned multifunctional radar countermeasures, a Partially Observable Markov Decision Model (POMDP) is used to model the confrontation game process between the electronic countermeasure aircraft and the target interference radar, where the confrontation game process is in a non-cooperative manner. The electronic countermeasure aircraft can only judge the working state of the radar by receiving the interference target radar signal and take corresponding tactical actions to block the "OODA" loop of the radar, so that it does not have the condition to transfer to other states.
[0053] In step S110, the multifunctional radar countermeasure process model is represented by a six-tuple (S, A, P, R, O, γ), which respectively includes: the working state space of the multifunctional radar, the interference action space, the state transition probability, the reward function, the observation space, and the discount factor.
[0054] Specifically, the working state space S of the multifunctional radar is a discrete space. According to the working process of a typical multifunctional radar seeker, when the seeker is continuously irradiated by interference, the radar may switch to the tracking interference source mode. The whole process is defined as seven states, that is, the working state space S includes 7 different working states, which are the coarse search working state, the precise search working state, the detection working state, the intercept working state, the range resolution working state, the stable tracking working state, and the passive tracking working state, denoted as s1, s2... s7. Among them, s 1~6 In ascending order of threat level, s7 only provides angle information, and its threat level is considered to be between the intercept working state s4 and the range resolution working state s5.
[0055] Furthermore, the belief state of the seeker is updated according to the Bayesian criterion based on the observed value:
[0056]
[0057] In the above formula, O(z, s′, a) is the observation function, representing the probability that the system is in state s′ when the observation value z is observed after performing action a; P(s′, s, a) is the state transition matrix, representing the probability of transitioning from state s to state s′ after performing action a; b(s) is the current belief state distribution, representing the probability that the system is in a certain state s. In each iteration, the system updates the belief state b′ based on the current belief state b(s) by performing action a and according to the new observation z.
[0058] As Figure 2 shown, it is a schematic diagram of the possible transition situations of the multi-functional radar under ECM (electronic jamming measures carried by electronic countermeasure aircraft).
[0059] Specifically, the interference action space A, that is, different ways to deal with different working states of the interference target radar, is a discrete space, which includes 6 interference patterns, namely: dense false target interference, Doppler scintillation interference, noise repetition interference, chaff interference, velocity range gate pull-off, and cross-eye interference, denoted as a1, a2... a6, and at the same time, turning off the interference is denoted as a7.
[0060] Furthermore, according to the principle of maximum entropy, assuming that in the absence of interference, the state transition probability of the seeker follows a uniform distribution. To evaluate the interference effect, an interference effectiveness factor δ (where 0 < δ < 1) is introduced, expressed as:
[0061]
[0062] In the above formula, the threat level of s′ is higher than that of s. δ can be used to measure the impact of interference on the state transition probability. When δ is close to 1, it indicates that the interference effect is small, and the state transition probability remains close to the uniform distribution without interference. When δ is close to 0, it indicates that the interference effect is large, and the state transition probability is significantly affected and may tend to certain specific state transitions. Therefore, the closer δ is to 1, the smaller the impact of interference on the seeker state transition. On the contrary, the closer δ is to 0, the greater the impact of interference on the state transition.
[0063] Specifically, the state transition probability P, that is, the probability of mutual transition between radar working states. After the multi-functional radar generates the radar task sequence, it will select radar phrases according to the target and environmental characteristics. Therefore, the strategy of radar phrase selection can be represented by p(P t |T t ,E t ), where E represents the target and environmental characteristics, and the state transition probability between radar states is expressed as:
[0064] p(S t+1 |St ) = P(S t+1 |T t )P(T t |S t )
[0065] In this embodiment, in order to enable the radar to transfer from a high-threat working mode to a low-threat working mode or a desired working mode in the shortest possible time, the reward function is designed as follows:
[0066]
[0067] r t = R base (s t+1 |s t , a t ) + R add (s t+1 |s t , a t )
[0068] In the above formula, T(s t ) represents the threat level when the multi-functional radar is in the state s t . The incentive value r t consists of two parts: the basic incentive R base and the additional incentive R add . If the interference causes the threat level of the radar working state to decrease, the basic incentive R base will give a positive incentive, otherwise a negative incentive. When the threat level of the radar working state is less than 0.1, it is considered that the threat of the multi-functional radar to the target is lifted, the interference is terminated, and an additional positive incentive will be given on the basis of the basic incentive, so as to ensure that the interference strategy can meet the ultimate goal of the game. This reward function can better guide the multi-functional radar to learn effective decision-making strategies in the reinforcement learning framework when dealing with complex electronic countermeasure (ECM) environments: 1) Through the combination of the basic incentive and the negative incentive, it helps to ensure that the radar acts immediately when it senses an increase in threat, thus effectively avoiding long-term exposure to a high-threat environment; 2) When the radar threat drops to a predetermined threshold (such as less than 0.1), an additional positive reward is given, which encourages the radar to lift the interference in the shortest possible time and restore or maintain the desired working mode; 3) Ensure that the radar not only focuses on short-term threat avoidance, but also can continuously optimize the long-term performance of the system after achieving short-term goals.
[0069] Specifically, the observation space O, that is, the characteristics of the radar emission signals of the interference targets detected by the electronic warfare aircraft. Since the working state of the multifunctional radar cannot be directly observed, it is necessary to interact with the multifunctional radar and predict the working state of the radar based on the characteristics of the received radar signals. The ESM system receives and measures the working parameters of the radar seeker to form radar phrases for identifying the working state. The observation space O is also the input data, and corresponding interference strategies are selected according to this data.
[0070] Therefore, in step S100, the characteristics of multiple radar signals and the corresponding working state labels are used as training data to train the interference strategy generation network.
[0071] Specifically, the characteristics of the radar signals are radar phrases formed by the working parameters of the multifunctional radar seeker.
[0072] Furthermore, assume that there are 15 radar phrases, denoted as z1, z2,..., z 15 Due to the fact that the ESM receiver has a false alarm probability (P f ) and an intercept probability (P d ) that cannot reach 100%, there is also a situation of empty observation, that is, no specific radar phrase is recognized, denoted as z 16 .
[0073] Assume that the observation probability satisfies a uniform distribution, and the observation probability is defined as follows: Except for s7, the probabilities of radiating z 16 in other states are the same, the probability of s7 radiating z 16 is 1 - P f , and the sum of the probabilities of the remaining observed values is P f . The observation probability function is expressed as:
[0074]
[0075] In the above formula, N(s i ) represents the number of possible observed values in each state of the seeker, and ∑P(z|s) = 1.
[0076] Specifically, a discount factor γ is used to measure the proportion of the current reward in the future total reward, and its value is γ ∈ (0, 1). It determines the trade-off relationship between the current reward and the future reward in the decision-making process: when γ is close to 1, the future reward is given a higher weight, and when γ is close to 0, more consideration is given to the current reward.
[0077] In this embodiment, according to the established multifunctional radar countermeasure process model, an A2C reinforcement learning network improved by using the attention mechanism is further established, that is, a Transformer network structure is used to fuse the A2C reinforcement learning method to construct an interference strategy generation network.
[0078] To solve the problems of long-term dependence and temporal credit assignment, in the A2C reinforcement learning network architecture, a Transformer network structure for spatio-temporal feature extraction is added. This structure includes multiple TIT units (TIT Blocks), and each TIT unit includes an Inner Transformer and an Outer Transformer.
[0079] Specifically, the Inner Transformer is responsible for processing the data of the current single observation, learning an effective observation representation to capture the important spatial information in the observation, while the Outer Transformer is responsible for processing multiple consecutive historical observations to capture the important temporal information across multiple observations. The two are combined to make better decisions. The two together form a TIT Block, and both spatial and temporal information can be fused in each TIT block; for stable training, through a dense connection design, all timely outputs of each TIT block are directed to the final output, so as to learn a better decision representation.
[0080] In this embodiment, before the data is input into the TIT Block, it also needs to be preprocessed. Given an observation array o ∈ R D each element in it is taken as a patch, that is: o p ∈ R N×1 , where N = D is the context length of the Encoder. Then, a trainable linear mapping E p is used to extract the Embedding feature, mapping each patch to a high-dimensional embedding space:
[0081]
[0082] In the above formula, E p ∈ R 1×DP , D p represents the dimension of the patch embedding feature.
[0083] At the same time, a trainable is added to capture the information of the entire sequence, expressed as:
[0084]
[0085] Since the Transformer network structure does not have the ability to capture the sequence order, a positional encoding is added to retain the position information, expressed as:
[0086]
[0087] Finally, use z0 as the input of the Inner Transformer in the first TIT Block.
[0088] In this embodiment, the interference strategy generation network is stacked by L TIT Blocks, and the operation of the nth TIT Block is expressed as:
[0089]
[0090] In the above formula, z l-1 represents the output data of the internal transformer of the previous TIT unit, represents the internal transformer, represents concatenating all z l [0] spanning K time steps, and L represents the number of TIT units. Therefore, spatio-temporal information can be fused in each TIT block, and this design enables the network to learn a more suitable representation form for decision-making tasks.
[0091] In this embodiment, a deep network is constructed by vertically stacking L TIT units to generate a hierarchical feature sequence:
[0092] <y1, y2, …, y L >
[0093] In the above formula, each y l contains the feature representation of K time steps, and D p is the feature projection dimension.
[0094] Furthermore, by densely connecting the last elements of all TIT units are concatenated, that is, is formed and the final policy π(·|o t ) is obtained through a feed-forward neural network and Softmax. Multiple stacked TIT units constitute the Enhanced_TIT backbone network for feature extraction.
[0095] Meanwhile, since both the Actor network and the Critic network take observations as inputs, the network parameters of multiple TIT units for feature extraction can be shared to accelerate network training. In this embodiment, the advantages of using a feedforward neural network (FFN) and Softmax to construct the Actor and Critic networks mainly lie in their high efficiency, stability, and applicability: the FFN enables fast computation and stable gradient propagation through a simple fully connected structure, and together with Softmax, it converts the Actor output into a probability distribution of discrete actions, naturally supporting the exploration-exploitation balance of the policy gradient algorithm; while the Critic, with the help of the nonlinear fitting ability of the FFN, provides reliable value estimates to guide policy optimization, and the two have clear divisions of labor and cooperate to improve the training efficiency.
[0096] In this embodiment, the reinforcement learning method A2C (Advantage Actor-Critic) algorithm is used to generate interference strategies under the above-mentioned multi-functional radar countermeasure process model. In A2C, the agent trains two networks simultaneously: the policy network (Actor) and the value network (Critic). The policy network is responsible for selecting the actions to be taken in a specific state, while the value network estimates the value of a given state (i.e., the expected future reward in that state). Different from traditional policy gradient methods, A2C improves the learning process by introducing the Advantage Function, reducing the variance in policy updates.
[0097] In A2C, the Advantage Function is defined as the advantage of choosing a certain action in the current state compared to the average policy. Specifically, the Advantage Function A(s,a) is calculated through the current state value function V(s) and the target return value:
[0098] A(s,a) = Q(s,a) - V(s)
[0099] In the above formula, Q(s,a) is the Q-value of taking action a in state s, and V(s) is the state value function of state s, representing the total return that the agent can expect to obtain starting from this state. The role of the Advantage Function is to tell the agent how much "advantage" a specific action has compared to other actions in a certain state.
[0100] In the process of policy optimization, A2C updates the policy network by using the policy gradient algorithm to maximize the expected cumulative return. The core idea of the policy gradient method is to adjust the policy parameters by calculating the gradient of the loss function, so that the agent takes the optimal action in a given state. In A2C, the calculation of the policy gradient is usually based on the current Advantage Function:
[0101] ▽ θ J(θ) = E t [▽θ logπ(a t |s t )(Q π (s t ,a t )-b t
[0102] In the above formula, J(θ) represents the performance of the target policy, and ▽ θ J(θ) represents the policy gradient, and π(a t |s t ) represents the probability of selecting action a t in state s t .
[0103] By parallelizing the training and introducing the advantage function, A2C can significantly improve the learning efficiency, reduce unnecessary fluctuations during training, and accelerate the convergence process. In addition, A2C can balance the trade-off between exploration and exploitation, thus ensuring that the agent can effectively explore the new policy space during training while avoiding getting stuck in local optimal solutions.
[0104] In this embodiment, according to the constructed action space A and state S, a deep reinforcement learning neural network integrating the Enhanced_TIT backbone network, i.e., the interference policy generation network, is established. When training the interference policy generation network, set the number of network training times N episode , and the number of training steps per round is N step . Initialize the network structure, with the L-layer Enhanced_TIT as the core feature extraction layer. Concatenate the last features y l [K - 1] ∈ R 1×Dp along the sequence dimension to obtain y ∈ R 1 ×Dp , and input it into the FFN. Generate the action probability distribution (Actor network) and state value estimation (Critic network) through Softmax. Initialize the radar environment state and observe the features, extract multi-level spatio-temporal features using Enhanced_TIT, and decode and generate action a t and its logarithmic probability log π (a t ,s t ) by combining the historical interaction relationship; execute the action, update the radar state according to the interaction environment, and calculate the reward R base ; update the belief state b t , calculate A(s t ,a t ); calculate ▽ θ J(θ) and optimize the Transformer-A2C network; perform radar state transition until the training converges.
[0105] Further, when training the interference strategy generation network, first set the number of network training times N episode , and the number of training steps per round is N step . At the beginning of each iteration, randomly select the initial radar working state s, generate the corresponding observation value o according to the radar state, initialize the belief state b, obtain the action a and logarithmic probability p from the Actor-Critic model, update the belief state to b' according to the observation value, and calculate the reward R base . If the threat level of the new observation state o is lower than the set threshold, then calculate R add , update the total reward r t , and update the Actor-Critic model, perform the radar state transition process, and update the radar state to s'. Repeat the training until the final condition is reached and the training is completed. Compared with the existing technology, it is creative in the update of the reward function and can optimize the state change of the radar in a complex interference environment.
[0106] Further, based on the total reward r obtained from the training t , evaluate the optimal interference decision for the multi-functional radar until the radar state is reduced to the target state, and obtain the interference strategy generation network based on reinforcement learning.
[0107] In step S130, after obtaining the trained interference strategy generation network, according to the transmitted signal of the interference target radar obtained in real time, input its features into the trained interference strategy generation network, and the interference strategy can be obtained in the multi-functional radar confrontation process model.
[0108] In this paper, the effectiveness of this method is also proved through simulation experiments. First, according to the above method, model the interference problem into a partially observable Markov decision process [S, A, P, R, O, γ], that is, the multi-functional radar confrontation process model. The elements in the model are described as follows: the multi-functional radar working state space S (s ∈ S), which includes 7 working states of the multi-functional radar. The interference pattern space A (a ∈ A), which includes 7 interference patterns that the interfering party can use. The state transition space P (p ij ∈ P), p ij is the probability that the multi-functional radar transfers from the working state s i to s j after interference, which defines the interaction relationship between the radar state and the interference. The reward function R includes R base and R add . Among them, R base is the basic incentive, which is used to measure the transfer of the radar working state from a high threat level to a low threat level or the desired working mode, and R base is the additional incentive. When the threat level of the radar working state is lower than 0.1, it will be in Rbase Give additional positive incentives on this basis. R can timely evaluate and feedback whether the radar threat level has decreased. The observation space O contains 16 observation values corresponding to the working states of the multifunctional radar. The discount factor γ measures the proportion of the current reward in the future total benefit, and γ ∈ (0, 1). Among them, the target interference radar working states, corresponding actions, and observation values in the simulation experiment are shown in Table 1.
[0109] Table 1 Radar states and corresponding actions and observation values
[0110]
[0111] Furthermore, use the reinforcement learning algorithm to train the neural network in the multifunctional radar countermeasure process model:
[0112] 1) Establish the initial Q-value table as follows:
[0113]
[0114] Among them, M is the number of multifunctional radar states, and N is the number of interferences.
[0115] 2) Radar interference strategy Transformer-A2C algorithm
[0116] Input: Radar state space S = [s1, s2,..., s7], interference pattern space A = [a1, a2,..., a7], Transformer encoding layer number L, historical observation window length H, discount factor γ = 0.99, learning factor l r = 0.001, exploration rate ε explore = 1 and ε explore_decay = 0.999.
[0117] Output: Trained Transformer-A2C network π θ * , optimal interference strategy.
[0118] Furthermore, when training the Transformer-A2C network:
[0119] 1. Model initialization
[0120] Construct a spatio-temporal feature extractor (Enhanced_TIT) and construct Actor and Critic models, and initialize the optimizer
[0121] 2. Reinforcement learning training loop (iterate N episodes times or until the termination condition is met):
[0122] a. Randomly select the initial radar state s i(i ∈ (1, 7))
[0123] b. Set the initial belief state b0 and construct a historical observation buffer
[0124] c. Loop N at each step step or until the termination condition is met: According to s i Generate the observation value o j (j ∈ [1, 16]), update the observation buffer by combining with historical observations; perform positional encoding on the observation sequence, generate multi-level spatio-temporal features y through Enhanced_TIT; with probability ε explore Randomly select an action, calculate the action log probability log π (a t |y); Execute the interference action a k (k ∈ [1, 7]), obtain the new state s i+1 and the reward r t ; Combine with the new observation value o j+1 , update the historical observation buffer; calculate the reward function R base , if T level ≤ 0.2, obtain R add ; Update the historical observation buffer; calculate the reward function R base , if T level ≤ 0.2, obtain R add ; Calculate the advantage function A t = R t - V(s t ) and update the Actor-Critic network parameters.
[0125] 3. Policy extraction
[0126] Freeze the Enhanced_TIT and Actor networks, extract the optimal interference policy a t = argmaxπ * (a|s t ).
[0127] As Figure 3 shown, it is the convergence schematic diagram after training using the A2C algorithm, indicating that as the number of training rounds increases, the low-threat ratio of the radar is increasing.
[0128] At the end of training, the optimal interference policy corresponding to each radar working state can be obtained, as shown in Table 2. However, it should be noted here that in the actual process, the working state of the adversarial target radar is still judged based on the current real-time observation value.
[0129] Table 2 Radar optimal interference policy
[0130]
[0131] In the above multi-functional radar jamming strategy generation method, comprehensively considering the confrontation between the electronic warfare aircraft and the multi-functional radar is a non-cooperative game process. The jamming decision-making problem is modeled as a partially observable Markov decision process, and a multi-functional radar threat assessment model based on the belief state space is introduced. By evaluating the threat, a revenue function for jamming decision-making, that is, a function is established. Finally, A2C is used to solve this problem to obtain the optimal jamming strategy. Compared with the existing technology, the jamming strategy method proposed by the method can effectively reduce the threat of the multi-functional radar to the target, has superior real-time performance and effectiveness, and significantly improves its practical application value.
[0132] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0133] In one embodiment, as Figure 4 shown, a multi-functional radar jamming strategy generation device is provided, including: a training data integration module 200, a jamming strategy generation network construction module 210, a jamming strategy generation network training module 220, and a jamming strategy real-time generation module 230, where:
[0134] The training data integration module 200 is used to obtain an external training data set, and the training data set includes the characteristics of various radar signals and the corresponding working state labels.
[0135] The non-cooperative game process modeling and jamming strategy generation network construction module 210 is used to adopt a partially observable Markov decision model to model the non-cooperative game process of multi-functional radar confrontation according to the "OODA" loop, obtain a multi-functional radar confrontation process model, and fuse Transformer and A2C networks to construct a jamming strategy generation network. The jamming strategy generation network includes a TIT feature extraction layer, a policy network for generating action probability distributions, and a value network for evaluating state values, thereby forming a double-branch decoupled structure.
[0136] The interference strategy generation network training module 220 is configured to train the interference strategy generation network using the training data set under the multi-functional radar countermeasure process model to obtain a trained interference strategy generation network.
[0137] The interference strategy real-time generation module 230 is configured to obtain the real-time radar signal emitted by the interference target and input the features of the real-time radar signal into the trained interference strategy generation network to generate a corresponding interference strategy.
[0138] For the specific limitations of the multi-functional radar interference strategy generation device, reference can be made to the limitations of the multi-functional radar interference strategy generation method described above, which will not be elaborated here.
[0139] Each module in the above multi-functional radar interference strategy generation device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0140] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a multi-functional radar interference strategy generation method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0141] Those skilled in the art can understand that Figure 5 the structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0142] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0143] Obtain a training data set, where the training data set includes the features of various radar signals and the corresponding working state labels;
[0144] Adopt a partially observable Markov decision model to model the non-cooperative game process of multi-functional radar countermeasure according to the "OODA" loop, obtain a multi-functional radar countermeasure process model, and fuse Transformer and A2C network to construct an interference strategy generation network. The interference strategy generation network includes a TIT feature extraction layer, a policy network for generating action probability distributions, and a value network for evaluating state values, thereby forming a double-branch decoupled structure;
[0145] Under the multi-functional radar countermeasure process model, use the training data set to train the interference strategy generation network to obtain a trained interference strategy generation network;
[0146] Obtain the real-time radar signal emitted by the interference target, and input the features of the real-time radar signal into the trained interference strategy generation network to generate a corresponding interference strategy.
[0147] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0148] Obtain a training data set, where the training data set includes the features of various radar signals and the corresponding working state labels;
[0149] Adopt a partially observable Markov decision model to model the non-cooperative game process of multi-functional radar countermeasure according to the "OODA" loop, obtain a multi-functional radar countermeasure process model, and fuse Transformer and A2C network to construct an interference strategy generation network. The interference strategy generation network includes a TIT feature extraction layer, a policy network for generating action probability distributions, and a value network for evaluating state values, thereby forming a double-branch decoupled structure;
[0150] Under the multi-functional radar countermeasure process model, use the training data set to train the interference strategy generation network to obtain a trained interference strategy generation network;
[0151] Obtain the real-time radar signal emitted by the interference target, and input the features of the real-time radar signal into the trained interference strategy generation network to generate a corresponding interference strategy.
[0152] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0153] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0154] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for generating a multi-functional radar jamming strategy, characterized in that, The method includes: Obtaining a training data set, which includes the features of various radar signals and the corresponding working state labels; Using a partially observable Markov decision model to model the non - cooperative game process of multifunctional radar countermeasure according to the "OODA" loop, obtaining a multifunctional radar countermeasure process model, and fusing Transformer and A2C network to construct an interference strategy generation network. The interference strategy generation network includes a TIT feature extraction layer, a policy network for generating action probability distributions, and a value network for evaluating state values, thus forming a double - branch decoupled structure; Under the multifunctional radar countermeasure process model, using the training data set to train the interference strategy generation network to obtain a trained interference strategy generation network; Obtaining the real - time radar signal emitted by the interference target and inputting the features of the real - time radar signal into the trained interference strategy generation network to generate a corresponding interference strategy.
2. The method for generating a multi-functional radar jamming strategy according to claim 1, wherein Each of the TIT units includes an internal transformer and an external transformer. In the l - th TIT unit: In the above formula, z l-1 represents the output data of the internal converter of the previous TIT unit, T l in represents the internal converter, represents concatenating all z l [0] over K time steps, and L represents the number of TIT units.
3. The method for generating a multi-functional radar interference strategy according to claim 2, wherein Both the policy network and the value network are constructed by longitudinally stacking multiple TIT units, a feed - forward neural network, and a Softmax layer.
4. The method for generating a multi-functional radar interference strategy according to claim 3, characterized in that, The multifunctional radar countermeasure process model includes: the working state space of the multifunctional radar, the interference action space, the state transition probability, the reward function, the observation space, and the discount factor.
5. The method for generating a multifunctional radar interference strategy according to claim 4, wherein The working state space includes: rough search working state, precise search working state, detection working state, intercept working state, range resolution working state, stable tracking working state, and passive tracking working state; The interference action space includes: dense false target interference, Doppler scintillation interference, noise repetition interference, chaff interference, velocity - range gate pulling, and cross - eye interference.
6. The method for generating a multi-functional radar jamming strategy according to any one of claims 1-5, characterized in that, The reward function is expressed as: r t = R base (s t+1 |s t , a t ) + R add (s t+1 |s t , a t ) In the above formula, T(s t ) represents the threat level when the multifunctional radar is in state s t , and the incentive value r t consists of a basic incentive R base and an additional incentive R add .
7. The method for generating a multi-functional radar interference strategy according to claim 6, wherein The feature of the radar signal is that a radar phrase is formed by the working parameters of the multifunctional radar seeker.
8. A multifunctional radar jamming strategy generation device, characterized in that The device includes: A training data integration module for obtaining external data to get a training data set, which includes the features of various radar signals and the corresponding working state labels; A non - cooperative game process modeling and interference strategy generation network construction module for using a partially observable Markov decision model to model the non - cooperative game process of multifunctional radar countermeasure according to the "OODA" loop, obtaining a multifunctional radar countermeasure process model, and fusing Transformer and A2C network to construct an interference strategy generation network. The interference strategy generation network includes a TIT feature extraction layer, a policy network for generating action probability distributions, and a value network for evaluating state values, thus forming a double - branch decoupled structure; An interference strategy generation network training module for training the interference strategy generation network using the training data set under the multifunctional radar countermeasure process model to obtain a trained interference strategy generation network; The interference strategy real-time generation module is used to obtain the real-time radar signal emitted by the interference target and input the characteristics of the real-time radar signal into the trained interference strategy generation network to generate the corresponding interference strategy.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Cognitive interference decision-making method based on deep reinforcement learning and application system thereof
CN121142484A
Complex mine radar adaptive anti-interference detection method based on reinforcement learning
CN121477158A