Frequency agility cognitive SAR anti-interference strategy generation method based on motion migration-demonstration learning
Through action transfer-demonstration learning and the D3QN algorithm, the interference countermeasure experience of the source environment is used to guide the generation of anti-interference strategies for the frequency-agile radar in the target environment. This solves the problem of slow strategy generation in the early stages of the confrontation, and improves the radar's anti-interference and imaging capabilities.
Patent Information
- Application Number
- CN202510806759.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
AI Technical Summary
In existing technologies, frequency-agile radars lack sufficient control over the battlefield in the early stages of a confrontation, and in modern electronic warfare environments, there is insufficient time to acquire a large number of adversarial samples, resulting in inefficient anti-interference strategy generation.
A method based on action transfer-demonstration learning is adopted. The interference confrontation experience in the source environment is used to pre-train the target environment. The D3QN algorithm is used to guide the learning of the optimal strategy in the target environment. A Markov decision process model is established, and a two-stage reward function is introduced to accelerate the generation of anti-interference strategies.
It realizes the rapid and accurate generation of anti-interference strategies for frequency-agile radar under different interference modes, and improves the anti-interference effect and imaging capability in the initial stage of confrontation.
Smart Images

Figure CN120706496A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of SAR anti-interference, and in particular relates to a SAR anti-interference strategy generation technology. Background Art
[0002] Synthetic Aperture Radar (SAR), a crucial military reconnaissance tool, can achieve high-resolution imaging in all weather conditions and at all times, making it an indispensable reconnaissance tool in modern electronic warfare.
[0003] However, in modern battlefield environments, with the development of jamming systems, particularly instantaneous frequency measurement technology, once the receiver in the jamming system acquires the SAR operating frequency, it can generate corresponding jamming signals, posing a serious threat to SAR imaging. Traditional anti-jamming methods are mainly divided into passive and active anti-jamming. Passive anti-jamming methods suppress interference through post-processing data, while active anti-jamming reduces the probability of interference by changing the parameters of the transmitted signal. Active anti-jamming is believed to provide more effective electronic countermeasures (ECCM) capabilities, and frequency agility (FA) is a commonly used active anti-jamming method. However, conventional frequency agility strategies cannot adapt to changes in the environment and jammers, making them difficult to cope with the increasingly complex jamming systems. Therefore, it is necessary to propose an intelligent generation method for FASAR anti-jamming strategies.
[0004] To ensure that radar's FA anti-jamming strategy can adapt to changes in the environment and jammers, some researchers have introduced reinforcement learning (RL) to design anti-jamming strategies. The paper "Reinforcement learning based anti-jamming frequency hopping strategies design for cognitive radar" (L. Kang, J. Bo, L. Hongwei, and L. Siyuan, IEEE Int. Conf. Signal Process., 2018) uses the signal-to-noise ratio (SINR) as a reward function and employs Q-learning and DQN to interact with the environment to learn FA radar anti-jamming strategies, successfully improving the detection signal-to-interference-and-noise ratio (SINR). The paper "Reinforcement learning based radar anti-jamming strategy design against anonstationary jammer" (J. Geng, B. Jiu, K. Li, Y. Zhao, and H. Liu, IEEE Int. Conf. Signal Process., 2022) uses detection probability as a reward to guide the generation of FA anti-jamming strategies. By combining RL and supervised learning, they enhance radar anti-jamming capabilities. The paper "An inverse reinforcement learning method to infer the reward function of an intelligent jammer" (Y. Fan, B. Jiu, W. Pu, K. Li, Y. Zhang, and H. Liu, Proc. IEEE Radar Conf., 2023) uses inverse reinforcement learning to infer the jammer's reward function based on its historical behavior and uses it to design a radar anti-jamming strategy. The paper "Transfer-Based DRL for Task Scheduling in Dynamic Environments for Cognitive Radar" (S. Akbar, R. S. Adve, Z. Ding and PWMoo, IEEE Transactions on Aerospace and Electronic Systems, vol. 60, no. 1, pp. 37-50, 2024) proposes a transfer reinforcement learning method that effectively improves the efficiency of cognitive radar task allocation.
[0005] The aforementioned RL-based FA radar anti-jamming strategy design methods all require a period of interaction between the radar and the jammer before generating an effective anti-jamming strategy. This reduces the frequency agile synthetic aperture radar (FASAR)'s ability to understand the battlefield in the early stages of a confrontation. Furthermore, in modern electronic warfare environments, there is often insufficient time to acquire a large number of adversarial examples for generating FASAR anti-jamming strategies, leading to their ineffectiveness. Transfer reinforcement learning, on the other hand, can leverage existing knowledge to accelerate training in new environments and enhance RL's adaptability in dynamic environments. Summary of the Invention
[0006] To solve the above technical problems, the present invention proposes a frequency-agile cognitive SAR anti-interference strategy generation method based on action transfer-demonstration learning, which effectively improves the anti-interference performance of FASAR.
[0007] The technical solution adopted by the present invention is: a method for generating frequency agile cognitive SAR anti-interference strategy based on action transfer-demonstration learning, comprising:
[0008] S1, using the frequency agile SAR signal for imaging;
[0009] S2. Establish a Markov decision process model for the SAR anti-interference strategy after frequency agility; the action space of the Markov decision process model is the transmit frequency of the SAR after frequency agility, the state space of the Markov decision process model is the detection state, transmit state and frequency of the jammer, the state transition function of the Markov decision process model is the transition probability of the interference state, and the reward function of the Markov decision process model includes the aperture level reward r a and pulse level rewards r p ;
[0010] S3. Anti-interference strategy generation method based on ATLFD. Specifically: consider two different but related adversarial environments, one denoted as the source environment and the other as the target environment. Use D3QN to solve the action value function, and divide the action value function estimation into two parts: the state value function V(s) and the advantage function A(s,a).
[0011] The interference adversarial experience in the source environment is used to pre-train the Markov decision process model in the target adversarial environment, and the expert network trained in the source environment is used to guide the learning of the optimal strategy in the target environment; thereby generating the optimal strategy.
[0012] Beneficial effects of the present invention: The method of the present invention utilizes transfer reinforcement learning to achieve rapid and accurate generation of FASAR anti-interference frequency agility strategies under different interference modes. The method of the present invention first models the confrontation problem between FASAR and the jammer as a Markov decision process, and proposes a two-stage reward shaping method to avoid the reward sparsity problem in the decision-making process. Then, an action transfer-demonstration learning (ATLfD) method is proposed, which pre-trains the FASAR in the target confrontation environment by utilizing the interference confrontation experience in the source environment, and uses the expert network trained in the source environment to guide the learning of the optimal strategy in the target environment. Simulation results show that compared with other existing transfer reinforcement learning methods, the method of the present invention effectively accelerates the generation of FASAR anti-interference strategies in the target confrontation environment, while improving the confrontation effect in the initial stage of the confrontation. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 This is a schematic diagram of the confrontation scenario between FASAR and jammer.
[0014] Figure 2 It is a schematic diagram of the transmission signal of FASAR and the detection / transmission mode of the jammer in different modes.
[0015] Figure 3 It is a schematic diagram of the multi-stage electronic countermeasure Markov decision process model.
[0016] Figure 4 It is the overall architecture of the ATLfD-based anti-interference strategy generation method.
[0017] Figure 5 is the average reward graph of different transfer reinforcement learning algorithms;
[0018] Among them, (a) is Figure 2 The average reward graph of different transfer reinforcement learning algorithms corresponding to interference mode 2 is (b) Figure 2 The average reward graph of different transfer reinforcement learning algorithms corresponding to interference pattern 3 in (c) is Figure 2 The average reward graph of different transfer reinforcement learning algorithms corresponding to interference pattern 4 in (d) is Figure 2 Average reward graphs for different transfer reinforcement learning algorithms corresponding to interference pattern 5.
[0019] Figure 6 is the average reward simulation graph for different reward shaping schemes;
[0020] Among them, (a) is Figure 2 The average reward simulation diagram of different reward shaping schemes corresponding to interference mode 1, (b) is Figure 2The average reward simulation diagram of different reward shaping schemes corresponding to interference mode 2, (c) is Figure 2 The average reward simulation diagram of different reward shaping schemes corresponding to interference mode 3, (d) is Figure 2 The average reward simulation diagram of different reward shaping schemes corresponding to interference mode 4, (e) is Figure 2 Simulation diagram of average rewards for different reward shaping schemes corresponding to interference mode 5.
[0021] Figure 7 It is the anti-interference strategy generated by D3QN and D3QN+ATLfD;
[0022] Among them, (a) is the anti-interference strategy corresponding to the first aperture time of D3QN, (b) is the anti-interference strategy corresponding to the first aperture time of D3QN+ATLfD, and (c) is the anti-interference strategy corresponding to the twentieth aperture time of D3QN+ATLfD. DETAILED DESCRIPTION
[0023] To facilitate those skilled in the art to understand the technical content of the present invention, the present invention is further explained below with reference to the accompanying drawings.
[0024] The present invention aims to address the slow generation of FASAR anti-interference strategies, which significantly reduces FASAR's anti-interference effectiveness in the initial stages of a confrontation. Therefore, the present invention proposes a solution for generating anti-interference strategies based on action transfer-demonstration learning. This method leverages demonstration data from the source environment and an expert action selection network to accelerate the generation of anti-interference strategies in the target environment.
[0025] Step 1: Establish a FASAR anti-interference model. Establish a FASAR and interference confrontation scenario such as Figure 1 As shown. Construct the interference signal model J(τ) and sort out the jammer detection / forwarding mode, as shown Figure 2 As shown in Figure 4, a FASAR anti-interference imaging model is further proposed.
[0026] Step 2: Establish a Markov decision process (MDP) to solve the adversarial scenario. Model the FASAR frequency agility strategy problem as a Markov decision process, such as Figure 3 As shown, the frequencies of the jammer and FASAR transmission signals are respectively used as state s t and action a t At the same time, a two-stage reward shaping is introduced to establish the pulse-level and aperture-level reward functions as the reward r in the Markov decision process. t . Convert the Markov decision process problem into a transfer reinforcement learning problem. Figure 3 The short dashed line in is used to indicate the start of the next state.
[0027] Step 3: Source environment training. In the source environment, the D3QN (Dueling Double Deep Q Network, deep reinforcement learning) algorithm is used to train the source environment's FASAR. The corresponding network parameters are And save the source environment adversarial data
[0028] Step 4: Pre-training the target environment. Using the source environment to fight against data And the loss function L D (θ), pre-training FASAR in the target environment, the corresponding network parameters are
[0029] Step 5: Target environment training. Use the target environment to fight against data To train FASAR in the target environment, the corresponding network parameters are At the same time, the action transfer module and the ε-AT greedy strategy are used to select the emission frequency a of the target environment FASAR t The overall algorithm process is as follows Figure 4 shown.
[0030] Step 6: Performance evaluation. Figure 2 The interference pattern 1 in is used as the source environment, and patterns 2-5 are used as the target environment for transfer reinforcement learning. Figure 2 In the jamming mode 1, the jammer will detect 1 FASAR pulse and transmit N j The jammer in jamming mode 2 will receive two FASAR pulses. When the two pulses received have the same frequency, the jammer will transmit N j The jammer in jamming mode 3 will receive two FASAR pulses. When the two pulses received have the same frequency, the jammer will transmit N j Otherwise, it will randomly select one frequency point from the two FASAR pulses it has received to transmit. The jammer in jammer mode 4 will receive three FASAR pulses. When the pulses it has received have more than or equal to two identical frequencies, the jammer will transmit N j The jammer in jamming mode 5 will detect 3 FASAR pulses, discard the first detected pulse, and transmit N pulses if the remaining two pulses have the same frequency. jOtherwise, a random transmission frequency will be selected from the remaining two FASAR pulse frequencies. j =5.
[0031] like Figure 5 Figure 2 shows the average rewards achieved by D3QN, D3QN+AT, D3QN+LfD, and the proposed D3QN+ATLfD method in adversarial situations. As can be seen, D3QN+ATLfD, proposed in this paper, has the fastest convergence speed under all interference modes and is even close to convergence at the very beginning of the adversarial situation. This is due to its full utilization of adversarial experience in the source environment and FASAR to guide the generation of anti-interference strategies, enabling rapid adaptation to new environments and optimal decision-making.
[0032] Figure 6 The figure shows the average rewards obtained by using different reward shaping methods in different interference modes, where AR represents the aperture-level reward, PR represents the pulse-level reward, and AP represents the aperture-level penalty. It can be seen from the figure that the PR+AP reward shaping method proposed in this paper has the best countermeasure effect in all interference modes. Except in interference mode 4, the final average normalized reward of PR+AP is close to 1, which means it has better stability than other reward shaping methods.
[0033] Figure 7 The FASAR transmission strategies obtained by D3QN and D3QN+ATLfD proposed in this invention in the first aperture time and the twentieth aperture time under interference modes 2-5 are shown. "Inter" in the figure represents that the jammer is in the detection state, and "f1~f3" represents the transmission of the corresponding frequency signals. Figure 7 It can be seen that the anti-interference strategy generated by D3QN+ATLfD proposed in the present invention within the first aperture time is already able to deceive the jammer's detection, while the anti-interference strategy generated by D3QN is difficult to play a role.
[0034] Through these steps, the present invention provides an effective method for quickly generating FASAR anti-interference strategies, which can effectively improve the confrontation performance of FASAR in the initial stage of confrontation and enhance the anti-interference imaging capability of FASAR. The flexibility and adaptability of this method can be applied to different interference modes and FASAR system configurations. The present invention verifies the performance of the proposed two-stage reward function and transfer reinforcement learning method through simulation experiments. The simulation results show that compared with the existing transfer reinforcement learning method, the action transfer-demonstration learning method designed by the present invention can significantly improve the anti-interference effect of FASAR in the initial stage of confrontation, ensuring the robust imaging of FASAR throughout the entire confrontation cycle.
[0035] The core solution of the present invention includes the following steps:
[0036] Step S1: Establishing FASAR anti-interference model
[0037] Assume that the jammer's operating mode alternates between receiving and transmitting. The jammer receives radar signals and transmits jamming signals of the corresponding frequency. The transmitted signal can be expressed as:
[0038] J(τ)=s j (τ)·exp(j2πf j τ) (1)
[0039] Where τ is time, f j is the frequency of the interference signal, s j (τ) is the baseband signal transmitted by the jammer. The present invention assumes that it is a broadband radio frequency interference, that is,
[0040] s j (τ)=α j exp(jπK j τ 2 ) (2)
[0041] Among them, α j and K j The present invention considers that the jammer has different detection / transmission modes.
[0042] Assuming that the FASAR platform is frequency agile in azimuth and can change the carrier frequency of each transmitted pulse, the FASAR echo is expressed as
[0043]
[0044] Where η is the slow time, ω r and ω a are the range and azimuth window functions respectively, R(η) represents the distance between the platform and the target, c is the speed of light, T a represents the synthetic aperture time, K r is the modulation frequency of the linear frequency modulation signal, and
[0045] f(η)=f c +△f η (4)
[0046] represents the FASAR signal carrier frequency that varies with azimuth and time, where f c is the center frequency, △f ηSince imaging with frequency agile SAR signals requires a series of complex data processing, the present invention considers FASAR to use only one frequency point data for sparse imaging (assuming that Figure 2 However, it should be emphasized that the present invention does not involve a specific sparse imaging method.
[0047] S2. Establishing a Markov decision process model for FASAR anti-interference strategy generation
[0048] The generation of cognitive FASAR anti-jamming strategy is essentially a multi-stage ECCM problem, which can be expressed as a Markov decision process (MDP) model. The interaction process between FASAR and jammer can be expressed as a four-tuple in They correspond to the FASAR transmission frequency, the jammer's detection / transmission state and frequency, the transition probability of the jamming state, and the instantaneous reward. The four-tuple is defined as follows:
[0049] S2.1 Action Space
[0050] According to formula (4), the set of FASAR optional carrier frequencies is expressed as
[0051] Ψ=f c +{△f,2△f,…K△f} (5)
[0052] Where K is the number of optional frequencies and ∆f is the frequency interval. Assume that the number of sampling points in the FASAR azimuth direction is T. The carrier frequency of each azimuth pulse can be expressed as an action in the FASAR anti-interference MDP problem, that is, Combined with the definition in formula (3), the action set of FASAR is Expressed as
[0053]
[0054] Where, is the frequency of the azimuth pulse in a given FASAR echo, t∈[1,T],k t ∈[1,K].
[0055] S2.2 State Space
[0056] According to the definition of jammer behavior in S1, the state set It can be expressed as
[0057]
[0058] in, is the carrier frequency of the interference signal emitted by the jammer in each azimuth direction, t∈[1,T],j t ∈[1,2], Inter represents the detection state, f i Indicates the interference frequency selected by the jammer. The detection status and interference frequency values under different interference modes are shown in Figure 2 , where the detection status specifically refers to the pulse status of the FASAR detected in this interference mode.
[0059] S2.3. State transfer function
[0060] In the cognitive FASAR anti-interference MDP problem, the state transition function represents that at time t, FASAR is based on the current interference state s t Take action t After that, the interference state at the next moment is s t+1 The probability p(s t+1 |s t ,a t The state transition functions of all states and actions can form a transition probability matrix P, which is directly related to the working mode of the jammer. The specific calculation method is:
[0061]
[0062] Among them, N i is the number of detection pulses of the jammer, N p It is the period of the jammer detecting and transmitting jamming signals.
[0063] S2.4 Reward Function
[0064] The present invention introduces a two-stage reward into the reward function, including the aperture-level reward r a and pulse level rewards r p .
[0065] Aperture level reward a Defined as:
[0066]
[0067] where t available It represents the number of available pulses within one aperture time, that is, the number of interference-free pulses with frequency f1.
[0068] Pulse Level Rewards p Defined as:
[0069]
[0070] Among them, B j represents the overlapped portion of the FASAR imaging signal and the interference signal frequency band, and B is the FASAR imaging signal bandwidth.p Indicates the degree of interference to the signal currently transmitted by FASAR.
[0071] In order to further improve the speed of FASAR anti-interference strategy generation, the aperture-level reward r a Rewritten as a penalty function:
[0072]
[0073] The introduction of aperture penalty can ensure that FASAR is aware of the defects of the current strategy at the end of each aperture and guide the learning direction of the optimal strategy. In summary, the two-stage reward function r t You can write:
[0074]
[0075] Among them, β a and β p are the coefficients of the two-stage reward function, and in the present invention, their values are 1 and 50 respectively. At time t, FASAR is based on the current interference state s t Select the frequency of the signal to be sent a t After that, you will receive a reward r for sending a certain frequency signal at the current moment t The goal of the FASAR anti-jamming problem is to find an optimal strategy π through continuous interaction with the jammer. * (a t |s t ), thereby maximizing the cumulative reward, which can be expressed as
[0076]
[0077] Among them, γ∈[0,1) is the discount factor, In order to obtain the mean value, in the present invention, the value is taken as 0.99.
[0078] S3. Anti-interference strategy generation method based on ATLFD
[0079] In order to solve the above-mentioned optimal anti-interference strategy generation problem, this paper proposes the ATLfD algorithm based on the transfer learning framework, which uses the knowledge of the source environment to accelerate the strategy learning speed in the target environment and realize the rapid generation of FASAR anti-interference strategy. Considering two adversarial environments with different jammer interference patterns, combined with the modeling method in step S2, the source adversarial environment model is represented as The target adversarial environment model is represented as and The difference lies in the different interference modes of the jammers in the environment, i.e. Figure 2 Different interference modes in, and their corresponding adversarial experiences are and
[0080] S3.1. Training in the source environment
[0081] The action value function represents the expected cumulative reward obtained by FASAR after taking action a in state s under a given strategy π. It can guide FASAR to take action a when the jammer state is s, and is expressed as
[0082]
[0083] When FASAR adopts the ε-greedy strategy, it will randomly select an action with probability ε, or select the best action with probability 1-ε In the present invention, the initial value of ε is set to 1, and in each round of confrontation, 5×10 -4 Continues to decay until it reaches a minimum value of 0.05.
[0084] In the present invention, D3QN is used to solve the action value function and complete the training in the source environment. During the training process, all the states of the jammer in this embodiment are based on Figure 2 The rules shown in the figure are generated (see step 6), and FASAR is based on the jammer state s at the current moment. t Use D3QN to select the action with the largest Q value as the launch action a at the current moment t , and obtain the corresponding reward r according to step S2.4 t , then the jammer transfers to the next state s according to the state transfer function in step S2.3 t+1 ,{a t ,s t ,r t ,s t+1} together constitute the training data set. D3QN divides the estimation of the action value function into two parts: the state value function V(s) and the advantage function A(s,a), which can be expressed as
[0085]
[0086] During each parameter update, V(s) is responsible for updating the baseline values of all FASAR actions under the current interference state, while A(s,a) is responsible for updating the specific values of the specific FASAR actions under that state. The D3QN method divides the trained network into a main network and a target network to avoid the overestimation error problem of traditional deep Q network methods. D3QN uses the main network and the target network to calculate the loss, namely:
[0087]
[0088] in, and Represents the parameter sets of the primary network and the target network in the source environment, is the experience buffer storage area in the source environment. Equation (15) represents the target value of the D3QN iteration in the source environment. At the same time, the main network in the source environment updates its parameters by minimizing (16) and copies these parameters to the target network at a lower frequency.
[0089] After training in the source environment, adversarial data in the source environment can be obtained and network parameters In the present invention, the number of training rounds is 30, and each round has 512 decision moments (each decision moment will complete the jammer to select the current jamming state s t FASAR selects action a t , FASAR received the award t , the jammer transfers to the next jamming state s t+1 This closed-loop process) network parameters will be updated after each decision moment.
[0090] S3.2. Pre-training in the target environment
[0091] Using source experience in a target environment with a different interference pattern than the source environment right The D3QN under is pre-trained, and its parameters are recorded as In this paper, pre-training consists of 1000 decision moments, and the network parameters are updated after each decision moment. The loss basis for network parameter update consists of four components: D3QN loss, marginal classification loss, multi-step temporal difference (TD) error loss, and L2 regularization loss, which can be written as:
[0092]
[0093] Where, is the D3QN loss shown in item (16), λ1, λ2, and λ3 are coefficients for balancing the losses, and their values are all 0.01. is the marginal classification loss, defined as
[0094]
[0095] in, and They represent the actions taken by the FASAR of the target environment and the FASAR of the source environment in state s, is a marginal function, when and When the same, the value is 0. and The value is 0.8 at different times.
[0096] The multi-step TD target is defined as
[0097]
[0098] We can further obtain the multi-step TD error loss function as
[0099]
[0100] The L2 regularization loss is:
[0101]
[0102] Among them, ||·||2 is the two-norm.
[0103] S3.3. Confrontation in the target environment
[0104] After the pre-training in the target environment is completed, the adversarial process in the target environment begins. During the adversarial period, The demo is initially copied to the target environment's replay buffer At the same time, new experiences generated in the target environment are also stored in When the buffer capacity is exceeded, the oldest adversarial experience will be overwritten in chronological order. In addition, when updating network parameters When , the loss is still calculated according to formula (18). If the current sampled data comes from the target environment, the marginal classification loss is not introduced, that is:
[0105]
[0106] At the same time, the present invention introduces an action transfer module in the target environment, using the FASAR trained in the source environment to guide the FASAR in the target environment to select an appropriate transmission frequency with a certain probability. The ε-AT greedy strategy is proposed to guide the FASAR's transmission frequency selection, namely:
[0107]
[0108] Among them, a random 、 Denote random actions, actions selected by the FASAR of the source environment, and actions selected by the FASAR of the target environment, respectively. ε The probability of selecting a random action and the action suggested by the source environment is used to balance the probability of selecting a random action. In this invention, it is set to 0.5, and p∈(0,1) is a random number. The input of the action migration module is the current jammer's transmission state, and the output is the action taken by the current FASAR.
[0109] S4. Get the optimal strategy π * (at |s t )
[0110] Those skilled in the art will appreciate that the embodiments described herein are intended to aid the reader in understanding the principles of the present invention, and it should be understood that the scope of the present invention is not limited to such specific descriptions and embodiments. Various modifications and variations are readily apparent to those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims.
Claims
1. A method for generating frequency-agile cognitive SAR anti-interference strategies based on action transfer-demonstration learning, characterized in that: include: S1, imaging using FASAR signals; S2. Establish a Markov decision process model for FASAR anti-interference strategy generation; the action space of the Markov decision process model is the FASAR transmission frequency, the state space of the Markov decision process model is the jammer's detection state, transmission state and frequency, the state transition function of the Markov decision process model is the transition probability of the interference state, and the reward function of the Markov decision process model includes the aperture level reward r a and pulse level rewards r p ; S3. Anti-interference strategy generation method based on ATLFD. Specifically: consider two different but related adversarial environments, one denoted as the source environment and the other denoted as the target environment. Both the source environment and the target environment include jammers and FASAR. The jammers in the source environment and the target environment use different interference modes. The action-value function represents the expected cumulative reward obtained by FASAR after taking action a in state s under a given strategy π; D3QN is used to solve the action-value function to obtain the optimal strategy.
2. The method for generating frequency agile cognitive SAR anti-interference strategy based on action transfer-demonstration learning according to claim 1, characterized in that: The specific process of using D3QN to solve the action-value function and obtain the optimal strategy is as follows: A1. Train D3QN in the source environment and obtain adversarial data in the source environment after training. and network parameters A2. Use Pre-train D3QN in the target environment, and obtain the network parameters in the target environment after training. A3. Start the confrontation process in the target environment. Copied to the target environment's replay buffer At the same time, the new experience generated in the target environment is stored in In the initial confrontation stage of the target environment, an action transfer module is introduced. The action transfer module uses the actions selected by FASAR given by the D3QN trained in the source environment to guide the actions selected by FASAR in the target environment.
3. The method for generating frequency agile cognitive SAR anti-interference strategy based on action transfer-demonstration learning according to claim 2, characterized in that: Step A1 specifically: D3QN divides the estimation of the action value function into two parts: the state value function V(s;θ) and the advantage function A(s,a;θ); the action value function is expressed as: in, Represents an action set The mean advantage function of all actions in , represents the size of the set, and θ is the network parameter; At the same time, the D3QN method divides the training network into a main network and a target network to avoid the overestimation error problem of the traditional deep Q network method. Therefore, the loss function used in the training process is: in, and Represents the parameter sets of the primary network and the target network in the source environment, It is a buffer storage area for experience in the source environment; the main network is minimized Update the parameters and copy them to the target network with a coefficient of 0.05; Complete training in the source environment; obtain adversarial data in the source environment and network parameters 4. The method for generating frequency agile cognitive SAR anti-interference strategy based on action transfer-demonstration learning according to claim 3, characterized in that: Step A2 is specifically as follows: In a target environment with a different interference pattern than the source environment, use Pre-train the D3QN in the target environment. The network parameters of the D3QN in the target environment are recorded as The loss function in the pre-training process consists of four parts: D3QN loss, marginal classification loss, multi-step temporal difference error loss, and L2 regularization loss, expressed as: Where, is the D3QN loss, is the marginal classification loss, is the multi-step time difference error loss, is the L2 regularization loss, and λ1, λ2, and λ3 are coefficients.
5. The method for generating frequency agile cognitive SAR anti-interference strategy based on action transfer-demonstration learning according to claim 4, characterized in that: In the adversarial process of step A3, if the current sampled data comes from the source environment, the loss function for updating the network parameters is the same as the loss function in the pre-training process.
6. The method for generating frequency agile cognitive SAR anti-interference strategy based on action transfer-demonstration learning according to claim 5, characterized in that: In step A3, during the adversarial process, if the current sampled data comes from the target environment, the loss function is expressed as:
7. The method for generating frequency agile cognitive SAR anti-interference strategy based on action transfer-demonstration learning according to claim 6, characterized in that: The action transfer module uses the FASAR selected actions given by the D3QN trained in the source environment to guide the selected actions of the FASAR in the target environment, which is specifically expressed as: Among them, a random 、 denote random actions, actions selected by FASAR of the source environment, and actions selected by FASAR of the target environment, respectively, and β ε is the coefficient, p is a random number, p∈(0,1).
Citation Information
Cited By
Satellite-borne cognitive SAR ground radio frequency interference suppression method and system
CN122307478A