Adaptive Intermittent Sampling and Repeater Jamming Waveform Design Method Based on JSR-PPO
By building a JSR-PPO network, using adaptive moment estimation Adam optimizer and real-time reward function, the problems of slow update of interference waveform parameters and poor generalization in the existing technology are solved, and efficient interference waveform design in complex radar environments are realized.
Patent Information
- Application Number
- CN202510596107.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The current technology has slow update speed of interference waveform parameters in complex radar echo environments, poor generalization of interference waveform parameters in deep reinforcement learning decisions, too relies on large sample data analysis, and distortion of interference effect evaluation methods in dynamic environments.
A JSR-PPO network connected by a policy subnet and a value subnet is constructed, and an adaptive moment estimation Adam optimizer is used to design the loss function of the importance sampling measurement function and the crop function. The real-time reward function of the JSR-PPO network is feedbacked and the interference effect is achieved through the real-time reward function of the JSR-PPO network to achieve end-to-end optimization.
It significantly improves the adaptability and real-time nature of the interference waveform in complex dynamic environments, enhances the generalization ability of unknown anti-radar echo environments, and improves interference stability and decision-making efficiency.
Smart Images

Figure CN120103280B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of radar technology, and further relates to a method for designing an adaptive intermittent sampling and forwarding interference waveform based on jamming-to-signal ratio-proximal policy optimization (JSR-PPO) in the field of electronic countermeasure technology. The present invention can be used to obtain an intermittent sampling and forwarding interference waveform that can counter a specific anti-jamming waveform through optimal decision-making according to different anti-jamming waveforms. Background Art
[0002] As the core of digital radio frequency memory (DRFM) technology, intermittent sampling and forwarding interference can generate coherent interference signals by intercepting radar pulse segments in real time and quickly forwarding them, and can form a multi-false target group with both suppression and deception effects at the radar receiving end. In the prior art, the interference waveform optimization method mainly relies on traditional optimization algorithms and reinforcement learning frameworks.
[0003] Zhang Ying disclosed a design method for an adaptive intermittent sampling and forwarding interference waveform in her published paper "Waveform Optimization Design of Intermittent Sampling and Forwarding Interference" (Master's Thesis of Harbin Engineering University, 2022. DOI: 10.27060 / d.cnki.ghbcu.2022.000858). This method optimizes waveform parameters through the Genetic Algorithm (GA) and the Deep Q-Network (DQN) algorithm respectively. The steps of the Genetic Algorithm (GA) adopted in this paper are as follows: First step, fix the radar signal length and the intermittent sampling and forwarding period, and calculate the number of intermittent sampling periods; Second step, randomly generate several groups of binary coding sequences as the number of initial populations; Third step, initialize the GA parameters, perform intermittent sampling and forwarding interference on each radar pulse, and calculate the ratio of the mean to the standard deviation of the interference signal amplitude after pulse compression; Fourth step, calculate the fitness function of each radar pulse; Fifth step, select, cross, mutate and update several population individuals to optimize the population; Sixth step, continuously loop the second to fifth steps until the population update converges to obtain the best gene coding sequence; Seventh step, decode the gene sequence to obtain the interference waveform corresponding to the best action set. The steps of the Deep Q-Network (DQN) algorithm adopted in this paper are as follows: First step, fix the radar signal length and the sampling period, and initialize the network parameters; Second step, randomly initialize a set of action sets as the initial strategy; Third step, according to the greedy strategy, select the action corresponding to the largest Q value in the Q table in each round of training as the action of this round; Fourth step, according to the selected action, perform intermittent sampling and forwarding interference on the radar signal, and calculate the reward value obtained according to the reward function; Fifth step, continuously repeat the third to fourth steps, and continuously update the network parameters according to the loss function until the loss function of the DQN network converges, and obtain the best action mapping as the best intermittent sampling and forwarding interference parameter. Although the two algorithms adopted in this paper can accurately train the waveform parameters of intermittent sampling and forwarding interference to a certain extent, there are still the following two deficiencies: First, when there are many received radar echoes and the dimensions are high, the optimization speed of GA and DQN in high-dimensional data is slow, resulting in poor convergence and stability of the optimization strategy of intermittent sampling and forwarding interference waveform parameters, and then causing the problem of low utilization efficiency of radar environment samples. This directly affects the optimization effect of the interference waveform, making the finally generated interference waveform may not be able to adapt to the rapidly changing electromagnetic environment, thus reducing the interference effect. Second, the core of using the DQN algorithm to optimize waveform parameters mainly relies on the experience replay mechanism. In a complex radar echo scenario, due to the highly dynamic and non-stationary characteristics of the signal environment, the historical data distribution stored in the traditional experience replay mechanism may not match the current environmental state, resulting in data distribution deviation during policy update, and then weakening the adaptability and robustness of the model to real-time echo characteristics.Therefore, when facing a complex and unknown electromagnetic environment, the interference waveform model trained by the DQN algorithm has weak generalization ability, poor optimization stability, convergence speed and generalization ability, and it is difficult to adapt to different radar waveforms and complex interference scenarios.
[0004] Sun Tao et al. disclosed an adaptive intermittent sampling and forwarding interference waveform design method based on reinforcement learning in their published paper "Adaptive Interference Waveform Design Method Based on Reinforcement Learning" (Aerospace Defense, 2021, 4(2): 59-66). The implementation steps of this method are as follows: First, obtain the radar echo signal and use it as the input of the interference waveform decision model; Second, set the constant false alarm rate CFAR (Constant False Alarm Rate) as the reward function of the Q-learning algorithm; Third, use the action-state table to select the action corresponding to the echo as the parameter of the interference signal, and calculate the reward value through the preset reward function; Fourth, repeat the first, second, and third steps until the action-state table converges, and use the action obtained at this time as the parameter of the interference signal. Although this method innovatively introduces the CFAR index as the reward function and solves the problem of the lack of a quantization benchmark for interference effects in traditional methods. However, this method still has two deficiencies: First, as an important evaluation index for the interaction between reinforcement learning and radar echoes, the reward function is directly related to the quality of the intermittent sampling interference waveform parameters selected. At the same time, the calculation of the CFAR probability depends on a large amount of historical echo data to construct a statistical reference benchmark. As a real-time interactive learning framework, Q-Learning cannot accumulate enough sample numbers in advance in dynamic confrontation, which will cause significant deviations in the estimated results of the reward value, directly leading to incorrect optimization of the intermittent sampling and forwarding interference waveform parameters and affecting the interference effect. Second, when the radar uses dynamic echoes, if the statistical characteristics of the historical echoes mismatch with the current echoes, it will further distort the evaluation of the CFAR reward function, causing the optimization direction of the intermittent sampling and forwarding interference parameters to deviate from the optimal solution and directly leading to a deterioration of the interference effect. For the above reasons, although this method can maintain basic functions in static scenarios, it is difficult to generate stable and effective interference waveforms under actual complex and variable radar echoes. Summary of the Invention
[0005] The object of the present invention is to provide an adaptive intermittent sampling and forwarding interference waveform design method based on JSR-PPO in view of the above-mentioned deficiencies of the prior art, aiming to solve the problems of slow update speed of interference waveform parameters in traditional methods when the radar echo environment is complex and has a high dimension, poor generalization of deep reinforcement learning decision interference waveform parameters when the radar environment changes dynamically, over-reliance on large sample data analysis, and distortion of the interference effect evaluation method in a dynamic environment.
[0006] The technical idea for achieving the object of the present invention is as follows: The present invention constructs a JSR-PPO network composed of a parallel connection of a policy sub-network and a value sub-network. Among them, the policy sub-network is composed of three fully connected layers and a normalization layer connected in series, and outputs a normalized probability distribution matrix; the value sub-network is composed of three fully connected layers connected in series, and outputs the value of the current policy. Through the policy sub-network therein, the intermittent sampling interference parameters are directly optimized in the high-dimensional continuous action space, avoiding the problem of slow optimization speed when the echo environment is complex and has a large number of dimensions. When training the JSR-PPO network of the present invention, an Adaptive Moment Estimation (Adam) optimizer is used to independently update the parameters of the policy sub-network and the value sub-network. By designing a loss function for the JSR-PPO network containing an importance sampling measurement function and a clipping function, the policy sub-network ensures optimization stability through the dual mechanisms of the importance sampling measurement function and the clipping function. The importance sampling measurement function calculates the action probability ratio of the new and old policies, quantifies the distribution difference before and after the policy update, and prevents policy mutation. The clipping function is used to forcibly limit the update step size to achieve end-to-end joint optimization, taking into account the flexibility of policy exploration and the robustness of the training process. Since the network training dynamically balances the relationship between exploration and exploitation, the time for each decision of the interference parameters is significantly shorter than that of the traditional method, solving the problem of poor applicability of the traditional method to complex and variable radar echo environments. The present invention innovatively constructs a real-time reward function based on the Jamming-to-Signal Ratio (JSR), directly calculates the instantaneous reward value using the energy ratio of the interference signal and the target echo within the current pulse period, and replaces the traditional constant false alarm probability evaluation mode that relies on offline statistics. Through the JSR-PPO joint closed-loop architecture, the present invention uses the JSR-PPO network framework, adds JSR as the network reward function therein, and uses this reward function to real-time feedback the relationship between the interference waveform and the radar echo. By real-time feedback of the interference effect by JSR, the problem of relying on a large amount of offline data in traditional interference effect evaluation is effectively solved. At the same time, when the radar echo is updated in real time, the interference effect can also be real-time feedback, successively solving the problem of real-time evaluation distortion of traditional interference effect evaluation methods.
[0007] The specific steps for achieving the object of the present invention are as follows:
[0008] Step 1: Establish a JSR-PPO network composed of a parallel connection of a policy sub-network and a value sub-network;
[0009] Step 2: Establish an interference training set;
[0010] Step 3: Design a real-time JSR-PPO network reward function based on the Jamming-to-Signal Ratio (JSR);
[0011] Step 4, design the total loss function of the JSR-PPO network containing the importance sampling measurement function and the clipping function;
[0012] Step 5, input the interference training set into the JSR-PPO network, adopt the Adam optimizer, and independently iteratively update the parameters of the policy sub-network and the value sub-network until the total loss function of the JSR-PPO network converges, and obtain the trained JSR-PPO network;
[0013] Step 6, use the trained JSR-PPO network to decide the optimal sampling times and forwarding times, and generate the intermittent sampling and forwarding interference signal.
[0014] Furthermore, the policy sub-network is composed of a first fully connected layer, a second fully connected layer, a third fully connected layer, and a normalization layer connected in series in turn. Among them, the first to third fully connected layers are all implemented by linear transformation and non-linear activation functions, and the normalization layer is implemented by the normalization exponential function softmax; set the input channel numbers of the first to third fully connected layers in the policy sub-network module to: 512, 128, 128; the output channel numbers are set to: 128, 128, 64.
[0015] Furthermore, the value sub-network is composed of a first fully connected layer, a second fully connected layer, and a third fully connected layer connected in series in turn; set the input channel numbers of the first to third fully connected layers in the value sub-network to: 512, 128, 128; the output channel numbers are set to: 128, 128, 1.
[0016] Furthermore, the interference training set is composed of at least three signals in a 1:1:1 ratio, where 1 is a linear frequency modulation signal LFM (Linear Frequency Modulation), 1 is a stepped frequency pulse signal SFP (Stepped Frequency Pulse Signal), and 1 is a chirp-frequency stepped frequency pulse signal CF-SFP (Chirp-Frequency Stepped Frequency Pulse Signal).
[0017] Furthermore, the real-time JSR-PPO network reward function based on the JSR of the signal-to-jamming ratio is as follows:
[0018] ;
[0019] Among them, represents the real-time reward value of the JSR-PPO network at the th moment, represents the power of the intermittent sampling and forwarding interference signal at the th moment, represents the The power of the radar echo signal received by the jammer at a certain moment, represents the power of the noise signal during the propagation of the radar echo signal at a certain moment.
[0020] Furthermore, the power of the signal is obtained by the following formula:
[0021] ;
[0022] where, represents the power value of the signal , = 0, 1,..., , represents the signal number of discrete sampling points, represents the summation operation, represents the absolute value operation.
[0023] Furthermore, the importance sampling measurement function is as follows:
[0024] ;
[0025] where, represents the importance sampling measurement function, represents the network parameters of the policy sub-network at the th moment, , respectively represent the importance sampling measurement probabilities at the th moment and the th moment, , respectively represent the actions at the th moment and the th moment, , respectively represent the states at the th moment and the th moment, represents the operation of a occurring under the condition of b.
[0026] Furthermore, the clipping function is as follows:
[0027] ;
[0028] where, represents the clipping function, represents the clipping interval, and its value range is ε ∈ [ 0 . 1 , 0 . 2 ] .
[0029] The total loss function of the JSR-PPO network is as follows:
[0030] ;
[0031] where, represents the total loss function of the JSR-PPO network, represents the loss function of the policy sub-network at the th moment, represents the loss function of the value sub-network at the th moment, represents the network parameters of the value sub-network at the th moment.
[0032] The loss function of the policy sub-network is as follows:
[0033] L t c l i p ( θ t ) = E [ m i n ( r t ( θ t ) A t , c l i p ( r t ( θ t ) , 1 − ε , 1 + ε ) A t ) ] ;
[0034] where, E [ ⋅ ] represents the empirical expectation operation, represents the minimum operation, represents the advantage function.
[0035] The loss function of the value sub-network is as follows:
[0036] L t V F ( ϕ t ) = E [ ( V ϕ t ( s t ) − V t t a r g e t ) 2 ] ;
[0037] where, represents the action-value estimation function at the th moment, represents the expected action value at the th moment.
[0038] Furthermore, the steps of using the trained JSR-PPO network to decide the optimal sampling times and forwarding times are as follows:
[0039] First, the jammer receives the echo signal transmitted by the radar, takes the echo signal received at the current moment as the current state and inputs it into the trained JSR-PPO network, and outputs the optimal action after decision;
[0040] Second, map the optimal action to the sampling times and forwarding times of the intermittent sampling and forwarding jamming.
[0041] Compared with the prior art, the present invention has the following advantages:
[0042] First, the present invention adopts the JSR-PPO network framework, and directly processes the optimization problem of high-dimensional interference parameters in the continuous action space through the policy gradient network in the PPO framework, overcoming the defects of low sample efficiency and unstable convergence of interference waveform parameter strategies existing in traditional methods under complex radar echo environment conditions, so that the present invention significantly improves the adaptability of the adaptive intermittent sampling interference waveform design in complex dynamic environments.
[0043] Second, when training the JSR-PPO network, the present invention uses the Adam optimizer to independently update the parameters of the policy sub-network and the value sub-network to achieve end-to-end joint optimization, taking into account the flexibility of policy exploration and the robustness of the training process. Through the importance sampling measurement function and the clipping function in the designed total loss function of the JSR-PPO network, the dynamic balance between the exploration and utilization of intermittent sampling and forwarding interference waveform parameters is achieved, overcoming the policy oscillation problem that easily occurs in the modeling of complex radar echo environments by traditional deep reinforcement learning methods, so that the present invention enhances the generalization ability of the intermittent sampling and forwarding interference waveform parameter decision-making model to unknown adversarial radar echo environments, enabling the intermittent sampling and forwarding interference waveform to adapt to more complex radar echo environments.
[0044] Third, the real-time JSR-PPO network reward function designed based on the JSR of the signal-to-jamming ratio, according to its characteristics of clear physical meaning and no dependence on historical echo data, breaks through the dependence of traditional interference effect evaluation methods on the analysis of offline radar echo data, so that the present invention realizes the instantaneous quantization and online feedback of the interference effect, effectively improving the real-time performance and evaluation accuracy of the intermittent sampling and forwarding interference waveform parameter decision-making.
[0045] Fourth, through the closed-loop optimization architecture of the signal-to-jamming ratio feedback and the JSR-PPO network, the present invention realizes the dynamic adaptive adjustment of the intermittent sampling and forwarding interference waveform parameters, overcoming the defect that it is difficult for the interference waveform parameter optimization in the existing interference effect evaluation methods to cope with the changing radar echo environment, so that the present invention significantly improves the interference stability and decision-making efficiency of the intermittent sampling and forwarding interference. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flowchart of the present invention;
[0047] Figure 2 is a schematic structural diagram of the network of the present invention;
[0048] Figure 3 is a simulation result diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0049] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0050] Refer toFigure 1 , a further detailed description of the specific implementation steps of the embodiments of the present invention is given.
[0051] The radar in the embodiments of the present invention is a pulse compression radar. There is a jammer in the radar detection area that emits intermittent sampling and forwarding jamming. There is one or more targets near the jammer, and there is a certain distance between the jammer and the leading target.
[0052] Step 1, establish a JSR-PPO network composed of a policy sub-network and a value sub-network in parallel, as Figure 2 shown.
[0053] Step 1.1, construct a policy sub-network composed of a first fully connected layer, a second fully connected layer, a third fully connected layer, and a normalization layer connected in series in sequence. The normalization layer is composed of a normalization exponential function softmax, and the fully connected layer is composed of a linear transformation and a non-linear activation function. As Figure 2 shown by the policy sub-network in the dashed box.
[0054] Step 1.2, construct a value sub-network composed of a first fully connected layer, a second fully connected layer, and a third fully connected layer connected in series in sequence, as Figure 2 shown by the value sub-network in the dashed box.
[0055] Set the network parameters of the JSR-PPO network as shown in Table 1:
[0056] Table 1
[0057]
[0058] Step 2, establish an interference training data set.
[0059] The interference training data set established in the embodiments of the present invention contains 3 signals, which are respectively composed of 1 linear frequency modulation signal (LFM), 1 stepped frequency pulse signal (SFP), and 1 chirp stepped pulse signal (CF-SFP). The three signals form a training set in a ratio of 1:1:1 to ensure the training balance and generalization ability of the JSR-PPO network.
[0060] Step 3, design a real-time JSR-PPO network reward function based on the jamming-to-signal ratio JSR.
[0061] The reward function is the core of reinforcement learning, which determines the optimization goal of reinforcement learning. Using the following formula, calculate the jamming-to-signal ratio at each moment and use the result as the reward value at that moment:
[0062] ;
[0063] where represents the The real-time reward value of the JSR-PPO network at a certain moment, denotes the power of the radar echo signal received by the jammer at the th moment, denotes the power of the intermittent sampling and forwarding interference signal at the th moment, denotes the power of the noise signal during the propagation of the radar echo signal at the th moment.
[0064] The power of the said signal is obtained by the following formula:
[0065] ;
[0066] where, denotes the power value of the signal , = 0, 1,..., , denotes the total number of discrete sampling points of the signal , denotes the summation operation, denotes the absolute value operation.
[0067] Step 4, design the loss function of the JSR-PPO network containing the importance sampling measure function and the clipping function as follows:
[0068] ;
[0069] where, denotes the importance sampling measure function, denotes the network parameters of the policy sub-network at the th moment, , respectively denote the importance sampling measure probabilities at the th moment and the th moment, , respectively denote the actions at the th moment and the th moment, , respectively denote the echo signals at the th moment and the th moment, denotes the operation of a occurring under the condition of b.
[0070] The said clipping function is as follows:
[0071] ;
[0072] where, Represents a clipping function, represents the clipping interval, and its value range is ε ∈ [ 0 . 1 , 0 . 2 ] .
[0073] The total loss function of the JSR-PPO network is as follows:
[0074] ;
[0075] Among them, represents the total loss function of the JSR-PPO network, represents the th moment's loss function of the policy sub-network, represents the th moment's loss function of the value sub-network, represents the th moment's network parameters of the value sub-network.
[0076] The loss function of the policy sub-network is as follows:
[0077] L t c l i p ( θ t ) = E [ m i n ( r t ( θ t ) A t , c l i p ( r t ( θ t ) , 1 − ε , 1 + ε ) A t ) ] ;
[0078] Among them, E [ ⋅ ] represents the empirical expectation operation, represents the minimum value operation, represents the advantage function.
[0079] The loss function of the value sub-network is as follows:
[0080] L t V F ( ϕ t ) = E [ ( V ϕ t ( s t ) − V t t a r g e t ) 2 ] ;
[0081] Among them, represents the th moment's action value estimation function, represents the th moment's expected action value.
[0082] Step 5: Train the JSR-PPO network to obtain a trained JSR-PPO network.
[0083] Step 5.1: Assume that the current moment is , randomly select a type of radar echo signal from the training set generated in Step 2 as the input signal of the JSR-PPO network , and map the echo signal to a normalized probability distribution matrix through the policy sub-network, and map the current input to its value through the value sub-network 。
[0084] Step 5.2, for the normalized probability distribution matrix perform weighted random sampling, that is, perform multiple independent random samplings according to the probability distribution to select the optimal action , and map the action to the sampling times and forwarding times of the intermittent sampling and forwarding interference at the current moment , and its mapping relationship is as follows: ;
[0085] ;
[0086] wherein, represents the sampling times at the current moment, represents the action selected at the current moment, represents the mapping coefficient, and in the embodiments of the present invention, this coefficient is selected but not limited to 8, represents the forwarding times at the current moment.
[0087] Step 5.3, generate the intermittent sampling and forwarding interference at the th time according to the sampling times and forwarding times obtained in Step 5.2.
[0088] Step 5.4, use the reward function in Step 3 to calculate the reward at the current moment . Combine the action selected in Step 5.2, its corresponding probability selected in the normalized probability distribution matrix , the reward at the current moment, and the input to form a quadruple , and store this quadruple in the experience pool of the JSR-PPO network.
[0089] Step 5.5, the policy sub-network ensures optimization stability through the dual mechanisms of the importance sampling measurement function and the clipping function. The importance sampling measurement function calculates the action probability ratio of the old and new policies, quantifies the distribution difference before and after the policy update, and prevents the policy from mutating. The clipping function is used to forcefully limit the update step size, that is, limit within the [ 1 − ε , 1 + ε ] interval, directly constrain the policy update step size, and avoid gradient anomalies. On this basis, the policy sub-network and the value sub-network are respectively based on the policy loss function and the value loss function Calculate the loss value, and independently perform gradient descent and backpropagation through the Adam optimizer to iteratively update their respective network parameters.
[0090] Step 5.6, repeat Steps 5.1 - 5.5 until the loss functions of the two sub - networks of JSR - PPO converge respectively, and the parameter updates of the two sub - networks are synchronized to the overall JSR - PPO network, then the training of the JSR - PPO network is completed. To achieve end - to - end joint optimization, taking into account the flexibility of policy exploration and the robustness of the training process.
[0091] Step 6, use the trained JSR - PPO network to decide the optimal sampling times and forwarding times, and generate intermittent sampling and forwarding interference signals.
[0092] Step 6.1, the jammer receives the echo signal transmitted by the radar, takes the echo signal received at the current moment as the current state and inputs it into the trained JSR - PPO network, and outputs the optimal action after decision - making.
[0093] Step 6.2, map the optimal action to the sampling times and forwarding times of intermittent sampling and forwarding interference.
[0094] The following further illustrates the effect of the present invention in combination with simulation experiments:
[0095] 1. Simulation experiment conditions.
[0096] The software platform for the simulation experiment of the present invention is: Windows 11 operating system and PyCharm2022.
[0097] In the simulation experiment of the present invention, the distance between the simulated radar and the jammer is 30 km, and false targets are placed 150 m and 200 m around the jammer respectively.
[0098] The interference training data set used in the simulation experiment of the present invention contains 3 signals, which are composed of 1 linear frequency - modulated signal (LFM), 1 stepped - frequency pulse signal (SFP) and 1 chirp - stepped - pulse signal (CF - SFP) respectively. The three signals form a training set in a ratio of 1:1:1. Among them, for the linear frequency - modulated signal used in the simulation experiment of the present invention, its pulse repetition time is 64 us, bandwidth is 16 MHz, and sampling frequency is 32 MHz. For the stepped - frequency pulse signal used, its pulse repetition time is 4 us, bandwidth is 16 MHz, and sampling frequency is 32 MHz. For the chirp - stepped - pulse signal used, its pulse repetition time is 4 us, bandwidth is 16 MHz, and sampling frequency is 32 MHz.
[0099] 2. Simulation content and result analysis.
[0100] The simulation experiment of the present invention is to use the method of the present invention and a method of the prior art to obtain the reward values under the same training set respectively, and then plot the change relationship between the obtained reward values and the training rounds into a curve as shown in Figure 3 shown below.
[0101] A method of the prior art adopted in the simulation of the present invention is as follows:
[0102] Zhang Ying proposed a method for designing an adaptive intermittent sampling and forwarding interference waveform based on DQN in her published paper "Waveform Optimization Design of Intermittent Sampling and Forwarding Interference" (Master's Thesis of Harbin Engineering University, 2022. DOI: 10.27060 / d.cnki.ghbcu.2022.000858).
[0103] The following further describes the effect of the present invention in combination with the Figure 3 simulation diagram.
[0104] Figure 3 The abscissa of Figure 3 is the total number of rounds of reinforcement learning training, with the unit of times, and the ordinate is the reward value corresponding to each round. The curve marked with circles in
[0105] represents the relationship curve between the training rounds and the reward obtained by simulating with the prior art method; the curve marked with squares represents the relationship curve between the training rounds and the reward obtained by simulating with the method proposed by the present invention. Figure 3 As shown in Figure 3 , the reward value obtained by the method of the present invention continuously increases with the increase of the training rounds. Under the same training set, the reward value obtained by the method of the present invention is always greater than the reward value obtained by simulating with the prior art method.
Claims
1. An adaptive intermittent sampling and forwarding interference waveform design method based on JSR-PPO, characterized in that, The specific steps of this waveform design method are as follows: Step 1: Establish a JSR-PPO network composed of a parallel connection of a policy sub-network and a value sub-network; Step 2: Establish an interference training set; Step 3: Design a real-time JSR-PPO network reward function based on the JSR of the signal-to-interference ratio; Step 4: Design the total loss function of the JSR-PPO network containing an importance sampling measurement function and a clipping function; Step 5: Input the interference training set into the JSR-PPO network, adopt the Adam optimizer, and independently update the parameters of the policy sub-network and the value sub-network until the total loss function of the JSR-PPO network converges, and obtain the trained JSR-PPO network; Step 6: Use the trained JSR-PPO network to decide the optimal sampling times and forwarding times, and generate an intermittent sampling and forwarding interference signal.
2. The adaptive intermittent sampling and forwarding interference waveform design method based on JSR-PPO according to claim 1, wherein The policy sub-network described in Step 1 is composed of a first fully connected layer, a second fully connected layer, a third fully connected layer, and a normalization layer connected in series in sequence. Among them, the first to third fully connected layers are all realized by a linear transformation and a non-linear activation function, and the normalization layer is realized by the normalization exponential function softmax; the input channel numbers of the first to third fully connected layers in the policy sub-network module are set to: 512, 128, 128; the output channel numbers are set to: 128, 128, 64.
3. The adaptive intermittent sampling and forwarding interference waveform design method based on JSR-PPO according to claim 1, wherein The value sub-network described in Step 1 is composed of a first fully connected layer, a second fully connected layer, and a third fully connected layer connected in series in sequence; the input channel numbers of the first to third fully connected layers are set to: 512, 128, 128; the output channel numbers are set to: 128, 128, 1.
4. The adaptive intermittent sampling and forwarding interference waveform design method based on JSR-PPO according to claim 1, characterized in that, The interference training set described in Step 2 is composed of at least three signals in a 1:1:1 ratio, where 1 is a linear frequency modulation signal LFM, 1 is a stepped frequency pulse signal SFP, and 1 is a chirp stepped pulse signal CF-SFP.
5. The adaptive intermittent sampling and forwarding interference waveform design method based on JSR-PPO according to claim 1, wherein The real-time JSR-PPO network reward function based on the JSR of the signal-to-interference ratio described in Step 3 is as follows: ; Among them, represents the real-time reward value of the JSR-PPO network at the th moment, represents the power of the intermittently sampled and forwarded interference signal at the th moment, represents the power of the radar echo signal received by the jammer at the th moment, represents the power of the noise signal during the propagation of the radar echo signal at the th moment.
6. The adaptive intermittent sampling and forwarding interference waveform design method based on JSR-PPO according to claim 5, characterized in that The power of the signal is obtained by the following formula: ; wherein, represents the power value of the signal , = 0, 1, ..., , represents the total number of discrete sampling points of the signal , represents the summation operation represents the absolute value operation.
7. The adaptive intermittent sampling and forwarding interference waveform design method based on JSR-PPO according to claim 1, wherein The importance sampling measurement function described in Step 4 is as follows: ; Among them, represents the importance sampling measurement function, represents the network parameters of the policy sub-network at the -th moment, , respectively represent the importance sampling measurement probabilities at the -th moment and the -th moment, , respectively represent the actions at the -th moment and the -th moment, , respectively represent the echo signals at the -th moment and the -th moment, represents the operation of a occurring under the condition of b.
8. The adaptive intermittent sampling and forwarding interference waveform design method based on JSR-PPO according to claim 7, wherein The clipping function described in Step 4 is as follows: ; Among them, represents a clipping function, represents a clipping interval, and its value range is .
9. The adaptive intermittent sampling and forwarding interference waveform design method based on JSR-PPO according to claim 8, characterized in that, The total loss function of the JSR-PPO network described in Step 4 is as follows: ; Among them, represents the total loss function of the JSR-PPO network, represents the loss function of the policy sub-network at the -th moment, represents the loss function of the value sub-network at the -th moment, represents the network parameters of the value sub-network at the -th moment; The loss function of the policy sub-network is as follows: ; Among them, represents the empirical expectation operation, represents the minimum value operation, represents the advantage function; The loss function of the value sub-network is as follows: ; Among them, represents the action value estimation function at the th moment, and represents the expected action value at the 10. The adaptive intermittent sampling and forwarding interference waveform design method based on JSR-PPO according to claim 1, characterized in that The steps of using the trained JSR-PPO network to decide the optimal sampling times and forwarding times described in Step 6 are as follows: The first step: The jammer receives the echo signal transmitted by the radar, takes the echo signal received at the current moment as the current state and inputs it into the trained JSR-PPO network, and outputs the optimal action after decision-making; The second step: Map the optimal action to the sampling times and forwarding times of the intermittent sampling and forwarding interference.
Citation Information
Patent Citations
Interference pattern and working parameter joint optimization method for multifunctional radar
CN116338599A
Radar main lobe forwarding interference resisting method based on reversible residual network
CN116540189A