Adaptive intermittent sampling and forwarding interference waveform design method based on JSR-PPO
Through the JSR-PPO network framework and real-time reward function, the problem of slow update speed and poor generalization of interference waveform parameters in complex radar echo environments is solved, and efficient and real-time interference waveform design and evaluation is achieved, which significantly improves interference stability and adaptability.
Patent Information
- Application Number
- CN202510596107.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The existing technology has slow update speed of interference waveform parameters in complex radar echo environments, poor generalization of deep reinforcement learning decisions, too relies on large sample data analysis, and distorted interference effect evaluation methods.
The JSR-PPO network framework is adopted to directly deal with the high-dimensional interference parameter optimization problem through parallel connection of policy subnet and value subnet, and a real-time reward function based on dry signal ratio is designed, and the dual mechanism of importance sampling and cropping function is used to achieve end-to-end joint optimization.
It significantly improves the adaptability of adaptive intermittent sampling interference waveform design in complex dynamic environments, enhances generalization capabilities, realizes real-time interference effect evaluation and dynamic adaptive adjustment, and improves interference stability and decision-making efficiency.
Smart Images

Figure CN120103280A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of radar technology, and further relates to an adaptive intermittent sampling and forwarding interference waveform design method based on Jamming-to-Signal Ratio-Proximal Policy Optimization (JSR-PPO) in the field of electronic countermeasure technology. The present invention can be used to obtain an intermittent sampling and forwarding interference waveform that can counteract the waveform through optimal decision-making according to different anti-interference waveforms. Background Art
[0002] Intermittent sampling and forwarding interference is the core of the digital radio frequency memory (DRFM) technology. By intercepting radar pulse fragments in real time and quickly forwarding them to generate coherent interference signals, multiple false target groups with both suppression and deception effects can be formed at the radar receiving end. In the existing technology, the interference waveform optimization method mainly relies on traditional optimization algorithms and reinforcement learning frameworks.
[0003] In his paper "Optimal Design of Waveform for Intermittent Sampling and Forwarding Interference" (Master's Thesis of Harbin Engineering University 2022.DOI:10.27060 / d.cnki.ghbcu.2022.000858), Zhang Ying disclosed a design method for adaptive intermittent sampling and forwarding interference waveform. This method optimizes waveform parameters through genetic algorithm GA (Genetic Algorithm) and deep Q learning algorithm DQN (Deep Q-Network). The steps of the genetic algorithm GA used in this paper are as follows: the first step is to fix the radar signal length and the intermittent sampling and forwarding period, and calculate the number of intermittent sampling periods; the second step is to randomly generate several groups of binary coding sequences as the initial population size; the third step is to initialize the GA parameters, perform intermittent sampling and forwarding interference on each radar pulse, and calculate the ratio of the mean and standard deviation of the interference signal amplitude after pulse compression; the fourth step is to calculate the fitness function of each radar pulse; the fifth step is to select, cross, mutate and update several population individuals to optimize the population; the sixth step is to continuously repeat the second to fifth steps until the population update converges and the optimal gene coding sequence is obtained; the seventh step is to decode the gene sequence to obtain the interference waveform corresponding to the optimal action set. The steps of the deep Q learning algorithm DQN used in this paper are as follows: the first step is to fix the radar signal length and sampling period and initialize the network parameters; the second step is to randomly initialize a set of action sets as the initial strategy; the third step is to select the action corresponding to the maximum Q value in the Q table in each round of training as the action of this round according to the greedy strategy; the fourth step is to perform intermittent sampling and forwarding interference on the radar signal according to the selected action, and calculate the reward value according to the reward function; the fifth step is to repeat the third and fourth steps continuously, and continuously update the network parameters according to the loss function until the loss function of the DQN network converges, and obtain the best action mapping as the best intermittent sampling and forwarding interference parameters. Although the two algorithms used in this paper can accurately train the intermittent sampling and forwarding interference waveform parameters, there are still the following two shortcomings: First, when the received radar echoes are large and the dimension is high, the optimization speed of GA and DQN is slow under high-dimensional data, resulting in poor convergence and stability of the intermittent sampling and forwarding interference waveform parameter optimization strategy, and then the problem of low utilization efficiency of radar environment samples. This directly affects the optimization effect of the interference waveform, making the final generated interference waveform unable to adapt to the rapidly changing electromagnetic environment, thereby reducing the interference effect. Second, the core of using the DQN algorithm to optimize waveform parameters mainly relies on the experience replay mechanism. In complex radar echo scenarios, due to the highly dynamic and non-stationary characteristics of the signal environment, the distribution of historical data stored in the traditional experience replay mechanism may not match the current environment state, resulting in data distribution offset when the strategy is updated, thereby weakening the model's adaptability and robustness to real-time echo characteristics.Therefore, when faced with a complex and unknown electromagnetic environment, the interference waveform model trained by the DQN algorithm has weak generalization ability, poor optimization stability, convergence speed and generalization ability, and is difficult to adapt to different radar waveforms and complex interference scenarios.
[0004] In their paper "Adaptive Interference Waveform Design Method Based on Reinforcement Learning" (Aerospace Defense, 2021, 4(2): 59-66), Sun Tao et al. disclosed a reinforcement learning-based adaptive intermittent sampling forwarding interference waveform design method. The implementation steps of this method are as follows: the first step is to obtain the radar echo signal and use it as the input of the interference waveform decision model; the second step is to set the constant false alarm probability CFAR (Constant False Alarm Rate) as the reward function of the Q-Learning algorithm; the third step is to use the action-state table to select the action corresponding to the echo as the parameter of the interference signal, and calculate the reward value through the preset reward function; the fourth step is to repeat the first, second, and third steps until the action-state table converges, and use the action obtained at this time as the parameter of the interference signal. Although this method innovatively introduces the CFAR indicator as a reward function, it solves the problem that the traditional method lacks a quantitative benchmark for interference effects. However, this method still has two shortcomings: First, the reward function, as an important evaluation index of the interaction between reinforcement learning and radar echo, is directly related to the quality of the selected intermittent sampling interference waveform parameters. At the same time, the CFAR probability calculation needs to rely on a large amount of historical echo data to build a statistical reference benchmark. Q-Learning, as a real-time interactive learning framework, cannot accumulate enough samples in advance in dynamic confrontation, which will cause significant deviations in the result of reward value estimation, which directly leads to errors in the optimization of intermittent sampling forwarding interference waveform parameters and affects the interference effect. Second, when the radar uses dynamic echoes, if the statistical characteristics of the historical echoes do not match the current echoes, the evaluation of the CFAR reward function will be further distorted, which will cause the optimization direction of the intermittent sampling forwarding interference parameters to deviate from the optimal solution, directly resulting in a worse interference effect. Based on the above reasons, although this method can maintain basic functions in static scenes, it is difficult to generate stable and effective interference waveforms under actual complex and changeable radar echoes. Summary of the invention
[0005] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and provide an adaptive intermittent sampling forwarding interference waveform design method based on JSR-PPO, aiming to solve the problems of slow updating of interference waveform parameters in traditional methods when the radar echo environment is complex and the dimension is high, poor generalization of interference waveform parameters in deep reinforcement learning decisions when the radar environment changes dynamically, over-reliance on large sample data analysis, and distortion of interference effect evaluation methods in dynamic environments.
[0006] The technical idea for achieving the purpose of the present invention is: the present invention constructs a JSR-PPO network composed of a strategy sub-network and a value sub-network in parallel, wherein the strategy sub-network includes three fully connected layers and a normalized layer connected in series, and outputs a normalized probability distribution matrix; the value sub-network includes three fully connected layers connected in series, and outputs the value of the current strategy. The intermittent sampling interference parameters are directly optimized end-to-end in the high-dimensional continuous action space through the strategy sub-network, avoiding the problem of slow optimization speed when the echo environment is complex and the dimension is large. When training the JSR-PPO network, the present invention adopts the Adaptive Moment Estimation Adam optimizer to independently update the parameters of the strategy sub-network and the value sub-network. By designing the loss function of the JSR-PPO network containing the importance sampling measurement function and the clipping function, the strategy sub-network ensures the optimization stability through the dual mechanism of the importance sampling measurement function and the clipping function. The importance sampling measurement function calculates the action probability ratio of the new and old strategies, quantifies the distribution difference before and after the strategy update, and prevents strategy mutation. The clipping function is used to force the update step size to achieve end-to-end joint optimization, taking into account the flexibility of strategy exploration and the robustness of the training process. Due to the dynamic balance between exploration and utilization in network training, the time for each decision on interference parameters is significantly shortened compared with the traditional method, which solves the problem that the traditional method is poorly applicable to complex and changeable radar echo environments. The real-time reward function based on the jamming-to-signal ratio (JSR) innovatively constructed by the present invention directly calculates the instantaneous reward value by using the energy ratio of the interference signal to the target echo in the current pulse period, replacing the traditional constant false alarm probability evaluation mode that relies on offline statistics. The present invention uses the JSR-PPO joint closed-loop architecture and the JSR-PPO network framework, adds JSR as a network reward function, and uses the reward function to provide real-time feedback on the relationship between the interference waveform and the radar echo. The real-time feedback of the interference effect by JSR effectively solves the problem of relying on a large amount of offline data in the traditional interference effect evaluation. At the same time, when the radar echo is updated in real time, the interference effect can also be fed back in real time, which solves the problem of real-time evaluation distortion of the traditional interference effect evaluation method.
[0007] The specific steps for achieving the purpose of the present invention are as follows:
[0008] Step 1: Establish a JSR-PPO network consisting of a strategy sub-network and a value sub-network in parallel;
[0009] Step 2, establish interference training set;
[0010] Step 3, design a real-time JSR-PPO network reward function based on the signal-to-dry ratio JSR;
[0011] Step 4, design the total loss function of the JSR-PPO network containing the importance sampling measurement function and the clipping function;
[0012] Step 5, input the interference training set into the JSR-PPO network, use the Adam optimizer, independently and iteratively update the parameters of the strategy sub-network and the value sub-network until the total loss function of the JSR-PPO network converges, and obtain the trained JSR-PPO network;
[0013] Step 6: Use the trained JSR-PPO network to determine the optimal number of sampling and forwarding times, and generate an intermittent sampling and forwarding interference signal.
[0014] Furthermore, the strategy subnetwork is composed of a first fully connected layer, a second fully connected layer, a third fully connected layer, and a normalization layer connected in series in sequence, wherein the first to third fully connected layers are all implemented by linear transformation and nonlinear activation functions, and the normalization layer is implemented by a normalized exponential function softmax; the number of input channels of the first to third fully connected layers in the strategy subnetwork module is set to: 512, 128, 128; the number of output channels is set to: 128, 128, 64.
[0015] Furthermore, the value sub-network is composed of a first fully connected layer, a second fully connected layer, and a third fully connected layer connected in series in sequence; the number of input channels of the first to third fully connected layers in the value sub-network is set to: 512, 128, 128; the number of output channels is set to: 128, 128, 1.
[0016] Furthermore, the interference training set consists of at least three signals in a 1:1:1 ratio, wherein one is a linear frequency modulation signal LFM (Linear Frequency Modulation), one is a stepped frequency pulse signal SFP (Stepped Frequency Pulse Signal), and one is a chirp-frequency stepped frequency pulse signal CF-SFP (Chirp-Frequency Stepped Frequency Pulse Signal).
[0017] Furthermore, the real-time JSR-PPO network reward function based on the signal-to-dry ratio JSR is as follows:
[0018] ;
[0019] in, Indicates The real-time reward value of the JSR-PPO network at the moment, Indicates The power of the interfering signal is intermittently sampled at each moment. Indicates The power of the radar echo signal received by the jammer at a certain moment, Indicates The power of the noise signal during the propagation of the radar echo signal at a certain moment.
[0020] Furthermore, the power of the signal is obtained by the following formula:
[0021] ;
[0022] in, Indicates signal The power value, =0,1,..., , Indicates signal The number of discrete sampling points, represents the sum operation, Indicates the absolute value operation.
[0023] Furthermore, the importance sampling measurement function is as follows:
[0024] ;
[0025] in, represents the importance sampling measure function, Indicates The network parameters of the strategy sub-network at each moment, , Respectively represent moment, The importance sampling of the moment measures the probability, , Respectively represent moment, The action of the moment, , Respectively represent moment, The state of a moment, Indicates that operation a occurs under the condition of b.
[0026] Furthermore, the clipping function is as follows:
[0027] ;
[0028] in, represents the clipping function, Indicates the clipping interval, and its value range is e ∈ [ 0 . 1 , 0 . 2 ] .
[0029] The total loss function of the JSR-PPO network is as follows:
[0030] ;
[0031] in, represents the total loss function of the JSR-PPO network, Indicates The loss function of the policy sub-network at each moment is: Indicates The loss function of the value sub-network at each moment is: Indicates The network parameters of the sub-network at each moment.
[0032] The loss function of the strategy subnetwork is as follows:
[0033] L t c l i p ( i t ) = E [ m i n ( r t ( i t ) A t , c l i p ( r t ( i t ) , 1 − e , 1 + e ) A t ) ] ;
[0034] in, E [ ⋅ ] represents the experience expectation operation, Indicates the minimum value operation. represents the advantage function.
[0035] The loss function of the value sub-network is as follows:
[0036] L t V F ( ϕ t ) = E [ ( V ϕ t ( s t ) − V t t a r g e t ) 2 ] ;
[0037] in, Indicates The action value estimation function at each moment is: Indicates The expected action value at that moment.
[0038] Furthermore, the steps of using the trained JSR-PPO network to determine the optimal number of sampling times and forwarding times are as follows:
[0039] In the first step, the jammer receives the echo signal transmitted by the radar, inputs the echo signal received at the current moment as the current state into the trained JSR-PPO network, and outputs the optimal action after decision-making;
[0040] In the second step, the optimal action is mapped to the sampling times and forwarding times of intermittent sampling and forwarding interference.
[0041] Compared with the prior art, the present invention has the following advantages:
[0042] First, the present invention adopts the JSR-PPO network framework and directly processes the high-dimensional interference parameter optimization problem in the continuous action space through the policy gradient network in the PPO framework, overcoming the defects of low sample efficiency and unstable convergence of interference waveform parameter strategy in complex radar echo environment conditions of traditional methods. The present invention significantly improves the adaptability of adaptive intermittent sampling interference waveform design in complex dynamic environments.
[0043] Second, when training the JSR-PPO network, the present invention uses the Adam optimizer to independently update the parameters of the strategy sub-network and the value sub-network to achieve end-to-end joint optimization, taking into account the flexibility of strategy exploration and the robustness of the training process. Through the importance sampling measurement function and the clipping function in the designed total loss function of the JSR-PPO network, the dynamic balance of the strategy exploration and utilization of the intermittent sampling forwarding interference waveform parameters is achieved, overcoming the strategy oscillation problem that is prone to occur in the traditional deep reinforcement learning method in the modeling of complex radar echo environments, so that the present invention enhances the generalization ability of the intermittent sampling forwarding interference waveform parameter decision model to unknown adversarial radar echo environments, so that the intermittent sampling forwarding interference waveform can adapt to more complex radar echo environments.
[0044] Third, the real-time JSR-PPO network reward function based on the interference-to-signal ratio JSR designed by the present invention has a clear physical meaning and does not need to rely on historical echo data. It breaks through the reliance of traditional interference effect evaluation methods on offline radar echo data analysis, enabling the present invention to achieve instantaneous quantification and online feedback of the interference effect, effectively improving the real-time performance and evaluation accuracy of intermittent sampling and forwarding interference waveform parameter decision-making.
[0045] Fourth, the present invention realizes dynamic adaptive adjustment of intermittent sampling and forwarding interference waveform parameters through the closed-loop optimization architecture of interference-to-signal ratio feedback and JSR-PPO network, overcoming the defect that the interference waveform parameter optimization in the existing interference effect evaluation method is difficult to cope with the changeable radar echo environment, so that the present invention significantly improves the interference stability and decision-making efficiency of intermittent sampling and forwarding interference. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flow chart of the present invention;
[0047] Figure 2 It is a schematic diagram of the structure of the network of the present invention;
[0048] Figure 3 It is a simulation result diagram of the present invention. DETAILED DESCRIPTION
[0049] The present invention is further described in detail below with reference to the figures and embodiments.
[0050] Reference Figure 1 , the specific implementation steps of the embodiment of the present invention are further described in detail.
[0051] The radar in the embodiment of the present invention is a pulse compression radar. A jammer is provided in the radar detection area to transmit intermittent sampling and forwarding interference. There are one or more targets near the jammer, and there is a certain distance between the jammer and the preceding target.
[0052] Step 1: Establish a JSR-PPO network consisting of a strategy sub-network and a value sub-network in parallel, such as Figure 2 shown.
[0053] Step 1.1, construct a strategy sub-network consisting of the first fully connected layer, the second fully connected layer, the third fully connected layer, and the normalization layer in series, where the normalization layer is composed of the normalized exponential function softmax, and the fully connected layer is composed of linear transformation and nonlinear activation function. Figure 2 The policy subnetwork is shown in the dashed box.
[0054] Step 1.2, construct a value sub-network consisting of the first fully connected layer, the second fully connected layer, and the third fully connected layer in series, such as Figure 2 The value subnetwork is shown in the dashed box.
[0055] The network parameters of the JSR-PPO network are set as shown in Table 1:
[0056] Table 1
[0057]
[0058] Step 2: Create an interference training data set.
[0059] The interference training data set established in the embodiment of the present invention includes three signals, namely, one linear frequency modulation signal (LFM), one stepped frequency pulse signal (SFP) and one frequency modulation stepped pulse signal (CF-SFP). The three signals constitute a training set in a 1:1:1 ratio to ensure the training balance and generalization ability of the JSR-PPO network.
[0060] Step 3, design a real-time JSR-PPO network reward function based on the dry-signal ratio JSR.
[0061] The reward function is the core of reinforcement learning. It determines the optimization goal of reinforcement learning. The following formula is used to calculate the signal-to-interference ratio at each moment, and the result is used as the reward value at that moment:
[0062] ;
[0063] in, Indicates The real-time reward value of the JSR-PPO network at this moment, Indicates The power of the radar echo signal received by the jammer at a certain moment, Indicates The power of the interfering signal is intermittently sampled at each moment. Indicates The power of the noise signal during the propagation of the radar echo signal at a certain moment.
[0064] The power of the signal is obtained by the following formula:
[0065] ;
[0066] in, Indicates signal The power value, =0,1,..., , Indicates signal The total number of discrete sampling points, represents the sum operation, Indicates the absolute value operation.
[0067] Step 4, design the loss function of the JSR-PPO network containing the importance sampling measurement function and the clipping function as follows:
[0068] ;
[0069] in, represents the importance sampling measure function, Indicates The network parameters of the strategy sub-network at each moment, , Respectively represent moment, The importance sampling of the moment measures the probability, , Respectively represent moment, The action of the moment, , Respectively represent moment, The echo signal at a certain moment, Indicates that operation a occurs under the condition of b.
[0070] The clipping function is as follows:
[0071] ;
[0072] in, represents the clipping function, Indicates the clipping interval, and its value range is e ∈ [ 0 . 1 , 0 . 2 ] .
[0073] The total loss function of the JSR-PPO network is as follows:
[0074] ;
[0075] in, represents the total loss function of the JSR-PPO network, Indicates The loss function of the policy sub-network at each moment is: Indicates The loss function of the value sub-network at each moment is: Indicates The network parameters of the sub-network at each moment.
[0076] The loss function of the strategy subnetwork is as follows:
[0077] L t c l i p ( i t ) = E [ m i n ( r t ( i t ) A t , c l i p ( r t ( i t ) , 1 − e , 1 + e ) A t ) ] ;
[0078] in, E [ ⋅ ] represents the experience expectation operation, Indicates the minimum value operation. represents the advantage function.
[0079] The loss function of the value sub-network is as follows:
[0080] L t V F ( ϕ t ) = E [ ( V ϕ t ( s t ) − V t t a r g e t ) 2 ] ;
[0081] in, Indicates The action value estimation function at each moment is: Indicates The expected value of an action at a given moment.
[0082] Step 5: Train the JSR-PPO network to obtain a trained JSR-PPO network.
[0083] Step 5.1, assuming the current time is , randomly select a type of radar echo signal from the training set generated in step 2 as the input signal of the JSR-PPO network , the echo signal is transformed into Mapped into a normalized probability distribution matrix , through the value sub-network, the current input The mapping of its value .
[0084] Step 5.2, normalize the probability distribution matrix Use weighted random sampling, that is, according to the probability distribution Perform multiple random independent sampling to select the best action , the action Mapped to the sampling times of intermittent sampling forwarding interference at the current moment and forwarding times , and its mapping relationship is as follows:
[0085] ;
[0086] in, Indicates the number of samples at the current moment. Indicates the action selected at the current moment. represents the mapping coefficient. In the embodiment of the present invention, this coefficient is selected but not limited to 8. Indicates the number of forwarding times at the current moment.
[0087] Step 5.3, according to the sampling times obtained in step 5.2 and forwarding times , generating Intermittent sampling and forwarding interference at each moment.
[0088] Step 5.4, use the reward function in step 3 to calculate the reward at the current moment . The action selected in step 5.2 , which is in the normalized probability distribution matrix Select The corresponding probability , the current reward , and input Together they form a quaternary , and store the quadruple into the experience pool of the JSR-PPO network.
[0089] Step 5.5: The policy subnetwork ensures optimization stability through the dual mechanisms of importance sampling measurement function and pruning function. The importance sampling measurement function calculates the action probability ratio of the new and old strategies, quantifies the distribution difference before and after the strategy update, and prevents strategy mutation. The pruning function is used to force the update step size to be limited. Restricted to [ 1 − e , 1 + e ] In the interval, the strategy update step is directly constrained to avoid gradient anomalies. On this basis, the strategy sub-network and the value sub-network are based on the strategy loss function And the value loss function The loss value is calculated, and the gradient descent and back propagation are performed independently through the Adam optimizer to iteratively update the respective network parameters.
[0090] Step 5.6, repeat steps 5.1 to 5.5 until the loss functions of the two sub-networks of JSR-PPO converge respectively, and the parameter updates of the two sub-networks are synchronized to the overall JSR-PPO network, then the training of the JSR-PPO network is completed. This achieves end-to-end joint optimization, taking into account the flexibility of strategy exploration and the robustness of the training process.
[0091] Step 6: Use the trained JSR-PPO network to determine the optimal number of sampling and forwarding times, and generate an intermittent sampling and forwarding interference signal.
[0092] In step 6.1, the jammer receives the echo signal transmitted by the radar, inputs the echo signal received at the current moment as the current state into the trained JSR-PPO network, and outputs the optimal action after decision making.
[0093] Step 6.2, mapping the optimal action to the sampling times and forwarding times of the intermittent sampling and forwarding interference.
[0094] The effect of the present invention is further described below in conjunction with simulation experiments:
[0095] 1. Simulation experimental conditions.
[0096] The software platform for the simulation experiment of the present invention is: Windows 11 operating system and PyCharm2022.
[0097] In the simulation experiment of the present invention, the distance between the simulated radar and the jammer is 30 km, and false targets are placed at 150m and 200m around the jammer respectively.
[0098] The interference training data set used in the simulation experiment of the present invention includes 3 signals, which are respectively composed of 1 linear frequency modulation signal (LFM), 1 step frequency pulse signal (SFP) and 1 frequency modulation step pulse signal (CF-SFP). The three signals constitute a training set with a 1:1:1 ratio. Among them, the linear frequency modulation signal used in the simulation experiment of the present invention has a pulse repetition time of 64us, a bandwidth of 16MHz, and a sampling frequency of 32MHz. The step frequency pulse signal used has a pulse repetition time of 4us, a bandwidth of 16MHz, and a sampling frequency of 32MHz. The frequency modulation step pulse signal used has a pulse repetition time of 4us, a bandwidth of 16MHz, and a sampling frequency of 32MHz.
[0099] 2. Simulation content and result analysis.
[0100] The simulation experiment of the present invention adopts the method of the present invention and a method of the prior art to obtain the reward values under the same training set respectively, and then plots the change relationship between the obtained reward value and the training round as shown in the figure: Figure 3 The curve shown.
[0101] A prior art method used in the simulation of the present invention is:
[0102] In his published paper "Waveform Optimization Design of Intermittent Sampling Forwarding Interference" (Master's Thesis of Harbin Engineering University 2022.DOI:10.27060 / d.cnki.ghbcu.2022.000858), Zhang Ying proposed an adaptive intermittent sampling forwarding interference waveform design method based on DQN.
[0103] Combine the following Figure 3 The simulation diagram of the present invention is further described.
[0104] Figure 3 The horizontal axis is the total number of reinforcement learning training rounds, and the vertical axis is the reward value corresponding to each round. Figure 3 The curve marked with a circle in the figure represents the relationship curve between the training rounds and the rewards obtained by simulating the prior art method; the curve marked with a square represents the relationship curve between the training rounds and the rewards obtained by simulating the method proposed in the present invention.
[0105] like Figure 3 As shown, the reward value obtained by the method of the present invention increases with the increase of the training round. Under the same training set, the reward value obtained by the method of the present invention is always greater than the reward value obtained by simulating using the prior art method.
Claims
1. A method for designing an adaptive intermittent sampling and forwarding interference waveform based on JSR-PPO, characterized in that: The specific steps of the waveform design method are as follows: Step 1: Establish a JSR-PPO network consisting of a strategy sub-network and a value sub-network in parallel; Step 2, establish interference training set; Step 3, design a real-time JSR-PPO network reward function based on the signal-to-dry ratio JSR; Step 4, design the total loss function of the JSR-PPO network containing the importance sampling measurement function and the clipping function; Step 5, input the interference training set into the JSR-PPO network, use the Adam optimizer to independently update the parameters of the strategy sub-network and the value sub-network until the total loss function of the JSR-PPO network converges, and obtain the trained JSR-PPO network; Step 6: Use the trained JSR-PPO network to determine the optimal number of sampling and forwarding times, and generate an intermittent sampling and forwarding interference signal.
2. The method for designing an adaptive intermittent sampling and forwarding interference waveform based on JSR-PPO according to claim 1 is characterized in that: The strategy subnetwork described in step 1 is composed of a first fully connected layer, a second fully connected layer, a third fully connected layer, and a normalized layer connected in series in sequence, wherein the first to third fully connected layers are implemented by linear transformation and nonlinear activation functions, and the normalized layer is implemented by a normalized exponential function softmax; the number of input channels of the first to third fully connected layers in the strategy subnetwork module is set to: 512, 128, 128; the number of output channels is set to: 128, 128, 64.
3. The method for designing an adaptive intermittent sampling and forwarding interference waveform based on JSR-PPO according to claim 1 is characterized in that: The value subnetwork described in step 1 is composed of the first fully connected layer, the second fully connected layer, and the third fully connected layer connected in series in sequence; the number of input channels of the first to third fully connected layers is set to: 512, 128, 128; the number of output channels is set to: 128, 128, 1.
4. The method for designing an adaptive intermittent sampling and forwarding interference waveform based on JSR-PPO according to claim 1 is characterized in that: The interference training set in step 2 is composed of at least three signals in a 1:1:1 ratio, wherein one is a linear frequency modulation signal LFM, one is a step frequency pulse signal SFP, and one is a frequency modulation step pulse signal CF-SFP.
5. The method for designing an adaptive intermittent sampling and forwarding interference waveform based on JSR-PPO according to claim 1 is characterized in that: The real-time JSR-PPO network reward function based on the signal-to-dry ratio JSR described in step 3 is as follows: ; in, Indicates The real-time reward value of the JSR-PPO network at this moment, Indicates The power of the interfering signal is intermittently sampled at each moment. Indicates The power of the radar echo signal received by the jammer at a certain moment, Indicates The power of the noise signal during the propagation of the radar echo signal at a certain moment.
6. The method for designing an adaptive intermittent sampling and forwarding interference waveform based on JSR-PPO according to claim 5 is characterized in that: The power of the signal is obtained by the following formula: ; in, Indicates signal The power value, =0,1,..., , Indicates signal The total number of discrete sampling points, represents the sum operation, Indicates the absolute value operation.
7. The method for designing an adaptive intermittent sampling and forwarding interference waveform based on JSR-PPO according to claim 1, characterized in that: The importance sampling measurement function described in step 4 is as follows: ; in, represents the importance sampling measure function, Indicates The network parameters of the strategy sub-network at each moment, , Respectively represent moment, The importance sampling of the moment measures the probability, , Respectively represent moment, The action of the moment, , Respectively represent moment, The echo signal at a certain moment, Indicates that operation a occurs under the condition of b.
8. The method for designing an adaptive intermittent sampling and forwarding interference waveform based on JSR-PPO according to claim 7 is characterized in that: The clipping function described in step 4 is as follows: ; in, represents the clipping function, Indicates the clipping interval, its value range is .
9. The method for designing an adaptive intermittent sampling and forwarding interference waveform based on JSR-PPO according to claim 8, characterized in that: The total loss function of the JSR-PPO network described in step 4 is as follows: ; in, represents the total loss function of the JSR-PPO network, Indicates The loss function of the policy sub-network at each moment is: Indicates The loss function of the value sub-network at each moment is: Indicates The network parameters of the value sub-network at each moment; The loss function of the strategy subnetwork is as follows: ; in, represents the experience expectation operation, Indicates the minimum value operation. represents the advantage function; The loss function of the value sub-network is as follows: ; in, Indicates The action value estimation function at each moment is: Indicates The expected action value at that moment.
10. The method for designing an adaptive intermittent sampling and forwarding interference waveform based on JSR-PPO according to claim 1, characterized in that: The steps for using the trained JSR-PPO network to determine the optimal number of sampling and forwarding times described in step 6 are as follows: In the first step, the jammer receives the echo signal transmitted by the radar, inputs the echo signal received at the current moment as the current state into the trained JSR-PPO network, and outputs the optimal action after decision-making; In the second step, the optimal action is mapped to the sampling times and forwarding times of intermittent sampling and forwarding interference.
Citation Information
Patent Citations
Interference pattern and working parameter joint optimization method for multifunctional radar
CN116338599A
Radar main lobe forwarding interference resisting method based on reversible residual network
CN116540189A
Intermittent sampling and forwarding interference suppression method based on Fast-SCNN
CN119375839A
Jamming signal generating apparatus and method thereof
US20220034997A1
Radar interference mitigation
US20220082654A1