Intermittent non-uniform sampling forwarding interference strategy optimization method, device and equipment

By optimizing radar jamming through intermittent non-uniform sampling and forwarding jamming strategy and parameterized action Markov decision process, the problems of easy identification of false targets and high complexity of deep reinforcement learning in existing technologies are solved, and efficient and reliable jamming strategy optimization and anti-jamming effect are achieved.

CN120703693APending Publication Date: 2025-09-26NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510816813.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

When facing modern anti-interference technology, the existing radar jamming technology generates false targets with strong regularity that are easy to identify and eliminate, and the deep reinforcement learning method has problems of limited algorithm scalability and high computational complexity when dealing with the optimization of intermittent sampling and forwarding jamming strategies.

Method used

An intermittent non-uniform sampling and forwarding jamming strategy is adopted. The jamming strategy is optimized through a parameterized action Markov decision process. The optimization objective is constructed by maximizing the mean absolute deviation of the reference unit power, sampling pulse width sequence and forwarding number sequence in the jamming round. The optimization objective is solved by the THPPO algorithm to achieve real-time response and multi-objective coordination of the jamming strategy.

Benefits of technology

It effectively improves the optimization efficiency and reliability of the jamming strategy, destroys the regularity of jamming, enhances the difficulty of anti-jamming, reduces the computational complexity, and realizes real-time response and multi-target coordination in confrontation scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120703693A_ABST
    Figure CN120703693A_ABST
Patent Text Reader

Abstract

The invention relates to an intermittent non-uniform sampling forwarding interference strategy optimization method, device and equipment. The method comprises the following steps: constructing an optimization target of an intermittent non-uniform sampling and forwarding interference confrontation process by maximizing average absolute deviation of reference unit power, a sampling pulse width sequence and a forwarding frequency sequence in an interference round, and modeling the intermittent non-uniform sampling and forwarding interference confrontation process as a parameterized action Markov decision process, wherein the state comprises actions of radar pulse width and historical time steps, the actions comprise sampling pulse width and corresponding forwarding times, the reward value comprises a target reward at the end of an interference round or a potential reward at the end of the interference round, an optimization target is reconstructed according to a parameterized action Markov decision process, and the reward value is obtained. And obtaining a reconstruction optimization target by maximizing an interference decision trajectory expectation accumulated reward sampled from the interference strategy, and solving the reconstruction optimization target to obtain an optimized interference strategy. By adopting the method, the optimization efficiency and reliability of the interference strategy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of radar technology, and in particular to a method, device and equipment for optimizing an intermittent non-uniform sampling and forwarding interference strategy. Background Art

[0002] With the rapid development of radar technology in various fields, including civilian traffic monitoring, the demand for efficient radar jamming and anti-jamming technologies has become increasingly urgent. Intermittent sampling and forwarding jamming (ISRJ), a coherent jamming technique, has become a common tool in radar countermeasures due to its fast response and low hardware resource consumption. Traditional ISRJ uses matched filtering at the receiver to generate coherent false targets to jam radar systems.

[0003] However, traditional ISRJ technology has significant drawbacks. The false targets it generates exhibit strong regularity in both range and amplitude dimensions, making them easily identified and eliminated by modern anti-jamming techniques. To overcome this limitation, numerous improvements have been proposed. For example, the intermittent sampling non-uniform periodic repetition (ISNPR) method, which disrupts radar detection by optimizing the non-uniform pulse width of jammer sub-pulses, and the non-periodic ISRJ (NP-ISRJ) technique, which significantly improves the jamming effectiveness against inverse synthetic aperture radar (ISAR). However, these methods often focus on optimizing the sampling pulse width and neglect adjusting the number of retransmissions. Some studies have used random sequences to modulate jamming waveform parameters to generate non-uniform, dense false target swarms, but this compromises jamming reliability. Methods that increase the number of false targets through clever noise convolution modulation require additional signal generators, significantly increasing system complexity and cost.

[0004] While deep reinforcement learning (DRL) technology has achieved some success in radar system optimization, it faces multiple technical bottlenecks when optimizing strategies for intermittent sampling and forwarding jamming (ISRJ). Existing DRL methods have inherent limitations when applied to parameterized action Markov decision processes (PAMDPs) with mixed action spaces. Traditional discretization or continuous approximation methods struggle to effectively adapt to mixed action spaces composed of continuous and discrete variables, limiting algorithm scalability. Furthermore, due to the curse of dimensionality and function approximation errors, the coupling relationship between action variables cannot be accurately characterized. At the model architecture level, specialized algorithms such as parameterized DQN (PDQN) and hybrid proximal policy optimization (HPPO) utilize fixed-dimensional input architectures. In dynamic scenarios, they rigidly compress multidimensional state information, forcing the policy network to process a large number of redundant state components. Such irrelevant inputs severely interfere with the policy selection mechanism, ultimately resulting in inefficient and unreliable jamming strategy optimization, making it difficult to meet the requirements of complex electromagnetic countermeasure scenarios. Summary of the Invention

[0005] Based on this, it is necessary to provide a method, device and equipment for optimizing intermittent non-uniform sampling and forwarding interference strategy to address the above technical problems.

[0006] A method for optimizing an intermittent non-uniform sampling and forwarding interference strategy, the method comprising:

[0007] The optimization objective of the intermittent non-uniform sampling and forwarding jamming countermeasure process is constructed by maximizing the mean absolute deviation of the reference unit power, sampling pulse width sequence, and forwarding number sequence in the jamming round.

[0008] The intermittent non-uniform sampling and forwarding interference countermeasure process is modeled as a parameterized action Markov decision process; each time step of the parameterized action Markov decision process includes the current state, action, state transition probability and reward value; the state includes the radar pulse width and the action of the historical time step; the action includes the sampling pulse width of the interference sub-pulse at the current time step and the corresponding number of forwarding; the state transition probability includes the probability of the jammer transferring from the current state to the next state after taking the interference action; the reward value includes the target reward at the end of the interference round or the potential reward before the end of the interference round; the potential reward includes the change in the interference signal distribution, the pulse width difference between the current sampling pulse width and the historical sampling pulse width, and the forwarding number difference between the current forwarding number and the historical forwarding number; the target reward includes the minimum average power of the reference unit, the average absolute deviation of the sampling pulse width and the forwarding number;

[0009] Reconstructing the optimization objective according to the parameterized action Markov decision process to maximize the expected cumulative reward of the interference decision trajectory obtained by sampling the interference strategy to obtain a reconstructed optimization objective; the interference decision trajectory includes a state sequence, an action sequence, and a reward sequence;

[0010] Solve the reconstruction optimization objective to obtain an optimized interference strategy.

[0011] An intermittent non-uniform sampling and forwarding interference strategy optimization device, the device comprising:

[0012] A target construction module is used to construct the optimization target of the intermittent non-uniform sampling and forwarding jamming countermeasure process by maximizing the mean absolute deviation of the reference unit power, sampling pulse width sequence, and forwarding number sequence in the jamming round;

[0013] A model building module is used to model the intermittent non-uniform sampling and forwarding interference countermeasure process as a parameterized action Markov decision process; each time step of the parameterized action Markov decision process includes the current state, action, state transition probability and reward value; the state includes the radar pulse width and the action of the historical time step; the action includes the sampling pulse width of the interference sub-pulse at the current time step and the corresponding number of forwarding; the state transition probability includes the probability of the jammer transferring from the current state to the next state after taking the interference action; the reward value includes the target reward at the end of the interference round or the potential reward before the end of the interference round; the potential reward includes the change in the interference signal distribution, the pulse width difference between the current sampling pulse width and the historical sampling pulse width, and the forwarding number difference between the current forwarding number and the historical forwarding number; the target reward includes the minimum average power of the reference unit, the average absolute deviation of the sampling pulse width and the forwarding number;

[0014] A target reconstruction module is used to reconstruct the optimization target according to the parameterized action Markov decision process to maximize the expected cumulative reward of the interference decision trajectory obtained by sampling the interference strategy to obtain a reconstructed optimization target; the interference decision trajectory includes a state sequence, an action sequence and a reward sequence;

[0015] The target solving module is used to solve the reconstruction optimization target and obtain the optimized interference strategy.

[0016] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0017] The optimization objective of the intermittent non-uniform sampling and forwarding jamming countermeasure process is constructed by maximizing the mean absolute deviation of the reference unit power, sampling pulse width sequence, and forwarding number sequence in the jamming round.

[0018] The intermittent non-uniform sampling and forwarding interference countermeasure process is modeled as a parameterized action Markov decision process; each time step of the parameterized action Markov decision process includes the current state, action, state transition probability and reward value; the state includes the radar pulse width and the action of the historical time step; the action includes the sampling pulse width of the interference sub-pulse at the current time step and the corresponding number of forwarding; the state transition probability includes the probability of the jammer transferring from the current state to the next state after taking the interference action; the reward value includes the target reward at the end of the interference round or the potential reward before the end of the interference round; the potential reward includes the change in the interference signal distribution, the pulse width difference between the current sampling pulse width and the historical sampling pulse width, and the forwarding number difference between the current forwarding number and the historical forwarding number; the target reward includes the minimum average power of the reference unit, the average absolute deviation of the sampling pulse width and the forwarding number;

[0019] Reconstructing the optimization objective according to the parameterized action Markov decision process to maximize the expected cumulative reward of the interference decision trajectory obtained by sampling the interference strategy to obtain a reconstructed optimization objective; the interference decision trajectory includes a state sequence, an action sequence, and a reward sequence;

[0020] Solve the reconstruction optimization objective to obtain an optimized interference strategy.

[0021] The above-mentioned intermittent non-uniform sampling and forwarding interference strategy optimization method, device and equipment, by constructing the optimization target by maximizing the average absolute deviation of the reference unit power, sampling pulse width sequence and forwarding number sequence in the interference round, can effectively break through the limitation of the traditional method of single-focus pulse width optimization, achieve the dual optimization of interference intensity improvement and interference regularity destruction, model the interference countermeasure process as a parameterized action Markov decision process, clarify the state, action, state transition probability and reward value of each time step, among which the state integrates the radar pulse repetition period and historical action, and the action covers the sampling pulse width and forwarding number, and clearly divides the action space into discrete (forwarding number) and continuous (sampling pulse width). The modeling of the wide) variables can achieve systematic decoupling of the decision space, significantly reducing the computational complexity brought by high-dimensional decision-making while ensuring the integrity of the optimization. The constructed parameterized action Markov decision process can accurately characterize the dynamic characteristics of the interference process and the mixed action space. The optimization target is reconstructed based on the parameterized action Markov decision process and converted into the expected cumulative reward for maximizing the interference decision trajectory. By designing a reward mechanism including potential rewards and target rewards, an adaptive coupling of electromagnetic environment conditions and interference parameters can be established. Potential rewards guide strategy exploration, and target rewards ensure long-term optimization direction, thus realizing real-time response and multi-target coordination of interference strategies. The embodiments of the present invention can effectively improve the optimization efficiency and reliability of interference strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 Schematic diagram of time-frequency modulation characteristics corresponding to INUSRJ in one embodiment;

[0023] Figure 2 1. A flow chart of a method for optimizing an intermittent non-uniform sampling and forwarding interference strategy according to an embodiment;

[0024] Figure 3 A dynamic diagram of state transition under the INUSRJ scheme in one embodiment;

[0025] Figure 4 A schematic diagram of a THPPO network architecture in one embodiment;

[0026] Figure 5 Schematic diagram of the comparison of the average round reward convergence trajectory of different DRL methods on each reward component in one embodiment, where (a) represents the total reward Rt Schematic diagram of the convergence characteristics, (b) represents the target reward r target Schematic diagram of the convergence characteristics, (c) represents the power regulation reward r power Schematic diagram of the convergence characteristics, (d) represents the interference parameter difference reward r MAD Schematic diagram of the convergence characteristics of ;

[0027] Figure 6 This is a structural block diagram of an apparatus for optimizing an intermittent non-uniform sampling and forwarding interference strategy in one embodiment;

[0028] Figure 7 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0030] This paper studies the electronic countermeasure scenario of a single jammer against a single-base radar system. The linear frequency modulation (LFM) signal emitted by the radar can be expressed as:

[0031]

[0032] Where, T represents the radar transmission pulse width, j is the imaginary unit, represents the FM slope, and B represents the corresponding FM bandwidth.

[0033] The jammer uses INUSRJ technology and uses K sampling pulses in each pulse repetition period (PRI). For the kth sampling pulse, τ k Indicates its sampling pulse width, z k The number of forwarding times is set. This configuration establishes the following sampling period relationship:

[0034] T s (k)=(z k +1)τ k (2)

[0035] Among them, T s (k) represents the duration of the kth interfering sub-pulse. Figure 1 The time-frequency modulation characteristic diagram corresponding to INUSRJ is shown in Figure 2. The intermittent non-uniform sampling pulse signal can be expressed as:

[0036]

[0037] in, It represents the time required for the first k-1 sampling pulses to complete the sampling and forwarding process, and its expression is:

[0038]

[0039] Combining formulas (1) and (3), we can get the complete mathematical expression of the sampled signal:

[0040]

[0041] By performing z on the kth sampling pulse k Repeated forwarding times, the complete mathematical model of the interference signal can be expressed as:

[0042]

[0043] Among them A j is the amplitude of the interfering signal.

[0044] The pulse pressure result of the z-th forwarding of the k-th sampling signal is:

[0045]

[0046] From the above formula, we can see that the main lobe center of the pulse pressure result of the kth sub-sampling signal forwarded for the zth time is located at The amplitude at this time is The main lobe width is Compared to classic ISRJ technology, INUSRJ demonstrates significant advantages through its non-uniform sampling and variable forwarding mechanism: different sampling pulses, combined with varying forwarding times, can generate false targets with varying amplitudes at different locations. The characteristics of these targets are closely related to the sampling width and forwarding times of each jammer pulse. Therefore, compared to ISRJ, the peak of the main lobe formed by the INUSRJ signal after pulse compression is generally lower. However, by properly setting the jamming parameters, a wider suppression range can be achieved. While achieving signal processing gain, the non-uniform sampling and forwarding mechanism effectively disrupts the regular characteristics of the false targets. This complex and variable signal form not only maintains the effectiveness of the jamming but also significantly increases the difficulty of anti-jamming countermeasures by increasing the unpredictability of the signal characteristics.

[0047] In one embodiment, Figure 2 As shown, a method for optimizing an intermittent non-uniform sampling forwarding interference strategy is provided, comprising the following steps:

[0048] Step 202 : constructing an optimization target for the intermittent non-uniform sampling and forwarding interference countermeasure process by maximizing the mean absolute deviation of the reference unit power, the sampling pulse width sequence, and the forwarding number sequence in the interference round.

[0049] To effectively interfere with target radar systems that use cell-averaged constant false alarm rate (CA-CFAR) detection, it is necessary to precisely configure the jamming signal parameters so that they fall into the reference cell, thereby increasing the radar detection threshold. Based on the working principle of INUSRJ technology, the spatiotemporal distribution characteristics of false targets formed by the jamming signal after pulse compression are mainly controlled by two key parameters: the sampling pulse width τ = [τ1,τ2,...,τ K ] and the number of forwarding z=[z1,z2,...,z K By jointly optimizing these two parameter vectors, we can simultaneously achieve: 1) enhancing the jamming signal's suppression of radar detection performance; and 2) increasing the complexity of jamming waveform feature recognition. This dual optimization mechanism improves jamming effectiveness while enhancing its ability to resist countermeasures.

[0050] Let x represent the radar waveform signal intercepted by the jammer. The jammer signal j(τ,z,p j ) corresponds to the time series representation of j(t) in equation (6), which is determined by three key parameters: sampling pulse width τ, forwarding times z and interference power p j Based on this, after the interference echo signal is processed by pulse compression, its output can be expressed as:

[0051]

[0052] The pulse compression process is implemented by simulating the radar receiver's signal processing chain from the jammer side using the intercepted radar signal. Therefore, the total power in the reference cell is calculated as follows:

[0053]

[0054] where variable n represents the number of reference units, ω n is a binary mask vector used to extract the values ​​from the reference cell set The mask vector is defined as follows:

[0055]

[0056] in, Indicates the reference unit index set corresponding to when the number of reference units is n.

[0057] Combining the above indicators, an objective function is constructed based on a dual mechanism: on the one hand, the reference unit power is increased to reduce the radar detection probability, and on the other hand, the regularity of the interference component is destroyed to hinder the extraction of interference component features. This objective function can be expressed as:

[0058]

[0059] Among them, d τ and d zRespectively represent the sampling pulse width sequence τ=[τ1,τ2,...,τ K ] and the forwarding number sequence z=[z1,z2,...,z K ] is the mean absolute deviation (MAD). As an important indicator for measuring the degree of dispersion of a data set, the introduction of the mean absolute deviation aims to enhance the differences between interference parameters, thereby effectively destroying the regularity of the interference signal. Compared with the variance, MAD shows stronger robustness when processing data sets containing outliers, which makes it particularly suitable for the application scenario of the present invention. Weight coefficient β τ and β z This reflects the relative importance of the differences between these two types of parameters. In addition, the present invention also imposes constraints on the sampling pulse width and the number of forwarding times.

[0060] Step 204: Model the intermittent non-uniform sampling and forwarding jamming countermeasure process as a parameterized action Markov decision process; each time step of the parameterized action Markov decision process includes the current state, action, state transition probability, and reward value; the state includes the radar pulse width and the action of the historical time step; the action includes the sampling pulse width of the jamming sub-pulse at the current time step and the corresponding number of forwarding; the state transition probability includes the probability of the jammer transitioning from the current state to the next state after taking the jamming action; the reward value includes the target reward at the end of the jamming round or the potential reward before the end of the jamming round; the potential reward includes the change in the jamming signal distribution, the pulse width difference between the current sampling pulse width and the historical sampling pulse width, and the forwarding number difference between the current forwarding number and the historical forwarding number; the target reward includes the minimum average power of the reference unit, the sampling pulse width, and the average absolute deviation of the forwarding number;

[0061] The optimization objective (11) involves mixed variables with complex coupling relationships, resulting in significant computational complexity. This challenge is further exacerbated in the multi-objective optimization process. The present invention is based on the parameterized action Markov decision process (PAMDP) framework to reconstruct the INUSRJ strategy optimization problem to address the limitations of traditional methods in dynamic adaptability and high-dimensional decision-making in mixed variable optimization. The PAMDP framework improves real-time response capabilities by establishing adaptive coupling between electromagnetic environment conditions and interference parameters, adopts a goal-oriented potential reward mechanism to alleviate the strategy deviation caused by multi-objective conflicts, and significantly reduces computational complexity while ensuring optimization integrity through systematic decoupling and reconstruction of the decision space.

[0062] The PAMDP framework extends the standard Markov decision process (MDP) to support mixed discrete-continuous action spaces. Based on this framework, the present invention proposes a hierarchical decoupling method for multidimensional decision spaces, which decomposes the complex decision process into a sequence of time-related sub-problems, each of which focuses on parameter optimization within a specific dimension, thereby significantly reducing the computational complexity. Specifically, the joint optimization of the sampling pulse width (τ) and the number of forwarding times (z) is mapped to a two-dimensional subspace through dimensionality reduction of the parameter space. At the decision execution level, the state value is constructed by combining historical trajectory data with real-time environmental parameters to capture environmental dynamics and enhance adaptability. Based on this framework, the actions generated at each decision stage are guided by the potential reward function, which effectively circumvents the preference convergence problem inherent in traditional optimization methods.

[0063] PAMDP usually consists of a five-tuple , where γ is the discount factor, which is used to balance the weights of immediate returns and future returns. The definitions and functions of the remaining elements are explained below.

[0064] The state refers to the jammer's observation of the environment information, such as Figure 3 The state transition dynamic diagram under the INUSRJ scheme shown in the figure integrates radar signal pulse width information and historical trajectory data. Its mathematical expression is:

[0065] s t =[T,τ1,z1,τ2,z2,...,τ t-1 ,z t-1 ] (12)

[0066] The state value integrates historical data from all interfering subpulses preceding the current subpulse. This integration enables the system to effectively utilize historical decision-making information to optimize future decisions, ensuring that each decision stage considers both past choices and future expectations. Furthermore, by incorporating real-time environmental parameters, the interference decision-making system achieves continuous environmental awareness, enabling dynamic adjustments to interference strategies based on environmental changes.

[0067] Action refers to the jammer's jamming parameters at the current moment, specifically the sampling pulse width τ of the t-th sub-pulse t (continuous action) and number of forwarding z t Joint decision making (of discrete actions). Hybrid action space The mathematical definition of is:

[0068]

[0069] Among them, the discrete action set Indicates the number of optional forwarding times of the current interference sub-pulse, continuous action set is related to z and represents the feasible sampling pulse width set of the current sub-pulse. Although there is a certain correlation between the discrete forwarding times and the continuous sampling pulse width, the present invention adopts a simplified assumption that the same sampling pulse width range is used under different forwarding times. Therefore, the continuous action set is uniformly expressed as The action space is simplified to:

[0070]

[0071] The interference action at the current moment can be expressed as a t =[z t ,τ t ].

[0072] The state transition dynamics is given by the conditional probability distribution P(s t+1 |s t ,a t ) determines the distribution defined in state s t Next, perform action a t Then transfer to state s t+1 The probability of . Figure 3 As shown, the present invention has a deterministic state transition characteristic. This determinism comes from the jammer's complete observability of historical parameter selection and environmental information. Its state transition dynamics can be formally expressed as:

[0073] s t+1 =Φ(s t ,a t )=[T,τ1,z1,τ2,z2,…,τ t ,z t ](15)

[0074] The reward function is used to quantify the immediate benefit of performing a specific action in a given state. When modeling the optimization of the interference strategy as a sequential decision-making process, the interference effect can only be evaluated after the complete strategy is executed. This temporal decoupling between the immediate parameter adjustment and the delayed evaluation of the interference effect can lead to the problem of sparse rewards. To effectively alleviate this problem, the present invention designs the potential reward r in a goal-oriented way. t int , which exploits the advantages of sequential decision making to provide better guidance for the learning and optimization of interference strategies.

[0075] The reward function is specifically expressed as follows:

[0076]

[0077] The parameter η t As the termination indicator of the current interference round. Specifically, when η t =1 indicates that the current interference round is completed, and the target reward value consistent with the original objective function (11) is generated. Used to evaluate interference performance; when η t = 0 means that the current interference round has not ended yet, and the parameters of the subsequent interference sub-pulses still need to be optimized. At this time, the potential reward value is generated The specific implementation details of the reward function are described as follows.

[0078] The potential reward value in formula (16) It is built on a goal-oriented mechanism and contains two key components: the previous state s t and the next state s t+1 The relative changes between the current selection parameters and the historical decision parameters. Its specific definition is:

[0079]

[0080] in, Indicates execution of action t The change in the distribution of the interference signal after For state s t The average amplitude of the interference signal distributed within the reference unit. By monitoring the changing trend of this indicator, the impact of parameter selection on the interference effect can be evaluated in real time. Represents the interference parameter τ executed at the current moment t With the historical interference parameter sequence [τ1,τ2,…,τ t-1 ] difference between. Parameter z representing the number of interference forwarding times executed at the current moment t With the historical interference parameter sequence [z1,z2,…,z t-1 ], where 1(·) is the indicator function. express The weight coefficient of the reward item, its core function is to adjust the reward index And the following two differentiated reward indicators D τ (s t ,a t )(pulse width difference) and D z (s t ,a t )(difference in the number of forwarding times). represents the target reward obtained after completing the decision of all sub-pulse interference parameters (determined by τ and z), which consists of the average power of the interference signal distributed in the reference unit and the mean absolute deviation (MAD) of τ and z. The target reward is consistent with the original target (11) and is defined as follows:

[0081]

[0082] in:

[0083] Represents the minimum average power of the interference signal distributed in the reference unit, where N is selected ref Typical reference cell configurations d τ and d z They are defined as the mean absolute deviation (MAD) of τ and z after removing the maximum and minimum values ​​corresponding to τ, respectively, to reduce the influence of extreme values. Indicates the granting of rewards The weight coefficient is designed to reduce with d τ d z The difference between the two rewards, while increasing their potential reward This weight configuration mechanism can effectively prevent the decision-making process from being overly influenced by potential rewards. , thus effectively preventing the final interference effectiveness from being adversely affected. τ and β z Used to constrain the impact of mean absolute deviation (MAD) on the overall reward: Among them, α corresponds to N ref The minimum threshold multiplication factor under different reference unit settings, p x Indicates the signal power of the detection unit where the real target is located, and are the corresponding weight coefficients respectively.

[0084] From β τ and β z It can be seen from the definition that only when More than p x This design achieves the optimal balance between interference effect and parameter diversity: First, ensure that the basic interference power meets the standard. This core operational requirement is met, and on this basis, through d τ and d z The reward value incentivizes maximizing parameter variability. This conditionally triggered dynamic optimization strategy fundamentally avoids the suboptimal decisions that can occur in traditional methods, where excessive pursuit of parameter variation sacrifices fundamental jamming effectiveness. This enables the jamming system to intelligently generate the optimal jamming waveform while ensuring effective suppression.

[0085] This paper models INUSRJ strategy optimization as an NP-hard mixed variable problem involving both discrete and continuous parameters, aiming to simultaneously maximize interference effectiveness and anti-interference difficulty. To overcome this mixed variable optimization challenge and break through the limitations of traditional methods, an innovative time-series state representation that integrates historical decision information and real-time electromagnetic environment parameters is introduced. Through context-aware evaluation and a dynamic environment-decision mapping mechanism, the limitations of traditional single-dimensional feedback are overcome, enabling continuous strategy updates without reinitialization. The designed goal-oriented potential reward mechanism effectively suppresses the bias-dominant phenomenon in high-dimensional greedy search algorithms, significantly reducing the risk of the algorithm falling into local optimality.

[0086] Step 206: reconstructing the optimization objective according to the parameterized action Markov decision process to maximize the expected cumulative reward of the interference decision trajectory obtained by sampling the interference strategy to obtain a reconstructed optimization objective; the interference decision trajectory includes a state sequence, an action sequence, and a reward sequence;

[0087] Under the framework of the parameterized action Markov decision process (PAMDP) proposed in this paper, combined with the designed reward function mechanism, the original optimization objective (11) is transformed into the following form:

[0088]

[0089] The trajectory l consists of the state, action and reward sequence generated by the current policy π, and its return represents the weighted sum of all reward values ​​on the trajectory. By optimizing the jamming strategy to maximize the expected cumulative reward throughout the entire jamming decision-making process, it aims to improve the radar detection threshold while destroying the regularity of the jamming component.

[0090] Step 208: Solve the reconstruction optimization target to obtain the optimized interference strategy.

[0091] In a specific embodiment of the present invention, a new deep reinforcement learning method based on the improved standard PPO framework, the THPPO algorithm, is proposed. Based on the standard PPO, the algorithm constructs a state temporal evolution model through the LSTM module to realize adaptive processing of variable-length state sequences, designs a dual-branch strategy network architecture to realize the decoupled generation and parallel optimization of discrete-continuous hybrid actions, and adopts a direct update mechanism in the strategy space to ensure learning stability.

[0092] The solution algorithm can also be a PPO framework embedded with an LSTM module, namely, the Time Series PPO (TPPO). When processing mixed action spaces, two special improved versions are used: discrete TPPO (dis-TPPO): the continuous action components are discretized into intervals of 500 and 1000 respectively; continuous TPPO (con-TPPO): discrete actions are modeled as continuous variables through a parameterized method.

[0093] In the above-mentioned intermittent non-uniform sampling and forwarding interference strategy optimization method, by constructing the optimization target by maximizing the average absolute deviation of the reference unit power, sampling pulse width sequence and forwarding number sequence in the interference round, it can effectively break through the limitation of the traditional method of single-focus pulse width optimization, and achieve the dual optimization of interference intensity improvement and interference regularity destruction. The interference countermeasure process is modeled as a parameterized action Markov decision process, and the state, action, state transition probability and reward value of each time step are clarified. The state integrates the radar pulse repetition period and historical action, and the action covers the sampling pulse width and forwarding number, and the action space is clearly divided into discrete (forwarding number) and continuous (sampling pulse width) variables. Separate modeling of quantities can achieve systematic decoupling of the decision space, significantly reducing the computational complexity brought by high-dimensional decision-making while ensuring optimization integrity. The constructed parameterized action Markov decision process can accurately characterize the dynamic characteristics and mixed action space of the interference process. Based on the parameterized action Markov decision process, the optimization objective is reconstructed and converted into the expected cumulative reward for maximizing the interference decision trajectory. By designing a reward mechanism including potential rewards and target rewards, an adaptive coupling of electromagnetic environment conditions and interference parameters can be established. Potential rewards guide strategy exploration, and target rewards ensure long-term optimization direction, thus realizing real-time response and multi-objective coordination of interference strategies. The embodiments of the present invention can effectively improve the optimization efficiency and reliability of interference strategies.

[0094] It is important to note that although existing PDQN and HPPO algorithms use a fixed-dimensional input architecture to process multidimensional state information, they exhibit limitations in dynamic state scenarios. This hard dimensionality constraint forces the policy network to process redundant state components. Such irrelevant inputs interfere with the policy selection mechanism, thereby reducing the efficiency and stability of the optimization process. The dual challenges of processing hybrid action spaces and representing dynamic states require innovative collaborative design of algorithmic architectures. An effective solution must integrate a dimensionally adaptive state encoding mechanism with a robust hybrid action processing framework to achieve simultaneous optimization of action selection accuracy and state representation.

[0095] Classical deep reinforcement learning algorithms are designed for discrete or continuous action spaces and are difficult to effectively handle the mixed action space studied in this paper. While discretization or continuous approximation methods can artificially unify the mixed space into a discrete or continuous space, these methods ignore the inherent structural hierarchical nature of the mixed action, leading to scalability bottlenecks caused by dimensionality explosion, approximation errors in constrained decision-making, and performance degradation due to loss of structural information. Furthermore, existing deep reinforcement learning frameworks have a fixed input dimension and are inefficient when processing states with dynamically changing dimensions. This fixed architecture forces the policy network to waste computational resources on irrelevant state features, which not only introduces noise into the gradient update process but also undermines convergence stability.

[0096] To overcome these limitations, we propose the THPPO algorithm, which is specifically optimized for hybrid action spaces and dynamic state processing. Experiments have shown that compared with traditional hybrid deep reinforcement learning methods (such as PDQN and HPPO), THPPO has significant advantages in training stability, time sequence processing capabilities, and scenario adaptability.

[0097] In one embodiment, solving the reconstruction optimization objective and obtaining the optimized interference strategy includes: constructing an interference decision model for predicting the action at each time step in each interference round; the interference decision model includes a dual-branch actor and a global critic, the dual-branch actor includes a discrete sub-actor and a continuous sub-actor, wherein the input layer of each network includes an LSTM layer, the global critic network is used to calculate the state value under the current state, and the advantage estimate calculated by using the state value is used to guide the strategy optimization of the discrete sub-actor network and the continuous sub-actor network, the discrete sub-actor network and the continuous sub-actor network are respectively used to select the number of forwardings and the sampling pulse width under the input state from the discrete action space and the continuous action space; the number of forwardings is obtained by sampling according to the probability distribution of the number of forwardings, and the sampling pulse width is obtained by matching the sampling pulse width Gaussian distribution parameters corresponding to the number of forwardings; the interference decision model is trained according to a preset loss function to obtain a trained dual-branch actor, and the real-time state is input into the trained dual-branch actor to obtain an optimized interference decision.

[0098] In this embodiment, the PAMDP framework decomposes high-dimensional interference parameter optimization into multiple sequential decision stages, each focusing on two key dimensions to achieve reduced computational complexity. The hierarchical dimensionality reduction transformation method proposed in this invention fundamentally reconstructs the original high-dimensional mixed action space optimization problem into a sequential decision-making process within a compressed two-dimensional subspace. This method reduces computational complexity by focusing each decision stage on a specific objective and parameter subset, thereby promoting efficient action space exploration. The instantaneous evaluation and feedback mechanism at each decision stage supports dynamic adjustment of subsequent decisions, effectively avoiding local optimality traps while improving optimization efficiency. By embedding historical decision trajectories into the current state representation, the temporal relevance of the decision process is maintained, providing context-aware evaluation capabilities for the entire decision sequence. Furthermore, the introduction of real-time environmental parameters significantly enhances the decision system's ability to adapt to dynamic environmental changes. Furthermore, because the optimization objective includes long-term cumulative rewards, the decision process fully considers the potential impact of current actions on future outcomes, exhibiting the characteristics of forward-looking planning. The resulting decision-making mechanism not only addresses the problem of high-dimensional complexity but also meets the adaptability requirements of dynamic environments.

[0099] To solve the established PAMDP problem, the present invention proposes the THPPO algorithm, which extends the standard PPO to support hybrid action space and dynamic state representation, thereby achieving robust policy updates in dynamically changing environments. The algorithm captures the temporal dependencies in variable-length state sequences by integrating LSTM modules, and adopts a parallel policy network architecture to jointly optimize discrete and continuous actions by directly updating in the policy space, further improving the stability of hybrid action learning. Experimental results show that compared with other deep reinforcement learning algorithms and optimization methods such as PSO and TS, the THPPO algorithm shows significant advantages in learning efficiency, stability, dynamic adaptability and generalization ability. The method of the present invention breaks through the limitations of the traditional static optimization paradigm by establishing a dynamic optimization mechanism that couples environmental perception with a closed loop of policy generation, and can achieve real-time adaptation to adversarial situations.

[0100] like Figure 4 As shown in the figure, a schematic diagram of the THPPO network architecture is provided. THPPO adopts an actor-critic architecture consisting of multiple parallel sub-actor networks and a global critic network. The mixed action space is decomposed into simpler action spaces. Each parallel sub-actor corresponds to a different action selection sub-problem. The present invention adopts two sub-actor networks, namely discrete sub-actor networks and and continuous sub-actor networks , and a global critic network as the state-value function To guide the update of the actor.

[0101] Since the above three networks need to process the state vector s containing historical parameters with temporal correlation characteristics t =[T,τ1,z1,τ2,z2,…,τ t-1 ,z t-1 ], which exhibits both sequential dependence and dimensional time-varying properties. To effectively handle such dynamically evolving states, all networks must be able to support variable-length input sequences. The proposed algorithm addresses this problem by embedding LSTM modules in the input layers of the three networks. This design not only explicitly models the temporal relationships between state parameters but also adaptively handles input sequences of arbitrary length.

[0102] Specifically, discrete sub-actors generate random strategies Input state value s t , output Z values ​​for Z possible discrete actions Further input it into the softmax layer to obtain the corresponding probabilities p1, p2, ..., p Z , and sample the current discrete action z t ,like Figure 4 As shown. Continuous child actor generation random strategy Similarly, the state s t As input, output the mean and variance of the corresponding continuous action, and sample from it to obtain the final continuous action value τ t Considering the continuous action τ t and discrete actions z t There is an internal connection between them, and they are not completely independent outputs. Therefore, the present invention outputs the conditional mean and variance parameters for each possible discrete action in the last layer of the continuous sub-actor network, which expands the output dimension of the continuous sub-actor network to Z times (the size of the discrete action space), thereby structurally coupling each discrete action selection with its associated continuous action distribution. The final continuous action τ t from With the selected z t Sampling is done from the corresponding distribution.

[0103] In one embodiment, the global critic includes an LSTM layer, multiple fully connected layers, and an output layer.

[0104] Specifically, the global critic is used to extract the temporal value features of the current state vector through the LSTM layer, process the temporal value features using multiple fully connected layers, and output the state value through the output layer.

[0105] In one embodiment, the discrete sub-actor includes an LSTM layer, multiple fully connected layers, a classification layer, and an output layer, and the continuous sub-actor includes an LSTM layer, multiple fully connected layers, and an output layer.

[0106] Specifically, in the discrete sub-actor, the LSTM layer is used to process the input state, extract the relationship between the state and discrete actions, and obtain discrete action features. Multiple fully connected layers are used to perform nonlinear transformations on the discrete action features to obtain enhanced discrete action features. The classification layer is used to process the enhanced discrete action features to obtain the probability distribution of discrete actions. The output layer is used to sample based on the probability distribution to obtain the number of forwarding times under the input state. In the continuous sub-actor, the input state is processed to extract the relationship between the state and continuous actions to obtain continuous action features. Multiple fully connected layers are used to perform nonlinear transformations on the continuous action features to obtain enhanced continuous action features. The output layer is used to obtain the Gaussian distribution parameters of the sampled pulse width corresponding to each possible number of forwarding times based on the enhanced continuous action features. The Gaussian distribution parameters of the sampled pulse width corresponding to the number of forwarding times output by the discrete sub-actor are then sampled to obtain the sampled pulse width under the input state.

[0107] In one embodiment, after the jammer makes an interference decision based on the trained interference decision model, the method further includes: performing time-width calibration on the sampling pulse width of the last interference sub-pulse of each interference round. In this embodiment, the present invention performs time-width calibration on the sampling pulse width of the last interference sub-pulse to ensure that the interference sampling signal can completely cover the radar pulse signal. This adjustment achieves a dual optimization effect: it ensures that the interference signal obtains pulse compression gain, and avoids using an excessively large sampling pulse width for the last sub-pulse. Although using an excessively large sampling pulse width for the last sub-pulse may greatly increase the mean absolute deviation (MAD) of the interference parameter, it has little effect on the improvement of the actual interference effectiveness.

[0108] In one embodiment, training an interference decision model based on a preset loss function to obtain a trained dual-branch actor includes: when the number of completed interference rounds reaches the model training frequency, updating the interference decision model based on sampled experience tuples and a preset loss function; the experience tuples include the state, state transition probability, reward value, and the number of forwarding times and sampling pulse width output by the interference decision model; and executing the next interference round using the updated interference decision model until the preset number of training times is met, thereby obtaining a trained dual-branch actor. In this embodiment, a direct update mechanism in the policy space is used to ensure learning stability. This enables continuous policy updates without the need for reinitialization.

[0109] In one embodiment, the method further includes: calculating the discounted reward corresponding to the state at each time step at the end of the interference round; the discounted reward is used to calculate the loss function of the global critic during model training; performing a generalized advantage estimation based on the state value, reward and state value of each time step and the next time step to obtain an advantage estimate; the advantage estimate is used to calculate the objective function of the discrete sub-actor and the continuous sub-actor during model training.

[0110] In one embodiment, updating the interference decision model based on the sampled experience tuples and a preset loss function includes: maximizing the loss function of the global critic based on the sampled experience tuples; the loss function of the global critic is obtained based on the mean square error of the state value and the discounted return; minimizing the objective functions of the discrete sub-actor and the continuous sub-actor based on the sampled experience tuples to update the interference decision model; the objective function of the discrete sub-actor is obtained based on the PPO clipping product of the probability ratio of the new and old strategies in the discrete action space and the advantage estimate, and the objective function of the continuous sub-actor is obtained based on the PPO clipping product of the probability density ratio of the new and old strategies in the continuous action space and the advantage estimate.

[0111] In this embodiment, a global critic network is used to estimate the state value function. Its design continues the standard PPO framework. This network uses state value prediction instead of state-action value prediction, avoiding the need to model Q(s) on a mixed action space. t ,a t ), thereby effectively alleviating the problem of parameter overestimation. The critic network is trained using the gradient descent method to optimize the parameters by minimizing the mean square error (MSE) between the predicted value and the empirical return. Its loss function is defined as:

[0112]

[0113] in Indicates that from state s t The starting experience reward estimate. To reduce the variance of the policy update The present invention introduces the advantage estimation function As a key indicator, this metric compares the action benefit with the baseline state value to quantify the relative advantages of actions. The generalized advantage estimation (GAE) method is used for analytical derivation, and its expression is:

[0114]

[0115] The THPPO algorithm uses the advantage estimate To update the sub-actor network. The discrete sub-actor network The objective function follows the clipping objective function of PPO, which ensures training stability by constraining the policy update amplitude. Its mathematical expression is:

[0116]

[0117] Where, Represents the ratio of two strategies Indicates that it is based on the old policy The new strategy obtained by updating is to use the clip(·) function to change the ratio Restricted to the interval [1-ε,1+ε] to ensure and The learning objective is to compare the clipped to unclipped strategies with an advantage estimate By taking the minimum of the clipped and unclipped targets, the algorithm suppresses the use of Updates beyond the clipping boundary. This mechanism achieves a balance between exploration and exploitation by allowing moderate updates to advantageous actions but restricting overly radical changes.

[0118] Continuous sub-actor network The objective function also follows the PPO clipping principle, and its mathematical expression is:

[0119]

[0120] Where, the continuous sub-actor network Output continuous action τ t Strategy Ratio is defined as:

[0121]

[0122] Where, and Represents the continuous action τ selected under the update strategy and the old strategy respectively t The THPPO algorithm ensures a stable and efficient learning process in the mixed action space by maintaining the same pruning mechanism, advantage estimation method and decoupled optimization architecture, thereby achieving effective updates of both discrete and continuous action policies.

[0123] In a specific embodiment, the training process of the THPPO algorithm is as follows:

[0124]

[0125] The proposed method is systematically verified through simulation experiments in radar-jammer confrontation scenarios. In order to rigorously evaluate the performance of the proposed framework, the present invention conducts comparative experiments with particle swarm optimization (PSO) method, TS algorithm and random parameter selection method. All algorithms follow the unified objective function defined by formula (11), and the parameter configuration is consistent with the final target reward defined by formula (18). The experiments were conducted under the same computational budget, using the number of interactions as a quantitative indicator. This indicator corresponds to the number of objective function evaluations for PSO / TS and the number of interference rounds for THPPO. In order to comprehensively evaluate the learning performance and optimization direction of each algorithm, the experimental results are presented from different reward component dimensions: target reward (r target ): The final goal defined in (18) Power regulation reward (r power ): Reference unit power control item Interference parameter diversity reward (r MAD ):(β τ d τ +β z d z ) quantifies the dispersion of interference parameters.

[0126] like Figure 5 The diagram shows a comparison of the average round reward convergence trajectory of different DRL methods on each reward component, where (a) represents the total reward R t Schematic diagram of the convergence characteristics, (b) represents the target reward r target Schematic diagram of the convergence characteristics, (c) represents the power regulation reward r power Schematic diagram of the convergence characteristics, (d) represents the interference parameter difference reward r MAD Schematic diagram of the convergence characteristics of Figure 5 (a) shows the total reward R tconvergence characteristics, among which the proposed THPPO algorithm shows better convergence robustness and obtains higher cumulative rewards than other algorithms. Compared with the TPPO algorithm, both versions of dis-TPPO show obvious volatility in the convergence process. This volatility is mainly due to the fact that the discretization of the action space cannot cover all potential effective actions, which is particularly prominent in tasks that require precise control, thereby limiting the effectiveness of policy learning and affecting the convergence stability. In addition, the analysis shows that the number of discrete intervals has a significant impact on the convergence performance of the dis-TPPO algorithm. On the other hand, the con-TPPO algorithm shows significantly lower learning efficiency, which is mainly attributed to the increased complexity of policy selection caused by the continuous transformation of the discrete action space. In the continuous action space, the agent needs to explore a wider range of actions, which essentially increases the complexity and duration of the search process. In addition, the policy update is often too conservative, and since each update is limited to a limited range, the learning process is further hindered. These fundamental limitations together show that continuous or discrete action space modeling has certain limitations in tasks that require both discrete selection and continuous control accuracy, thereby verifying the necessity of the hybrid action space decoupling mechanism adopted by THPPO. As Figure 5 As shown in (b), each algorithm has a target reward The convergence characteristics of the comparison results are compared with Figure 5 (a) and the analysis conclusions are consistent, further verifying the effectiveness of THPPO. Figure 5 (c)) and interference parameter diversification ( Figure 5 A comparative analysis of the two algorithms (d) reveals that while most algorithms can achieve baseline performance, THPPO demonstrates a significant advantage in simultaneously optimizing these two competing objectives. This performance advantage stems from three key design features employed by THPPO: a decoupled policy architecture, LSTM-enhanced temporal modeling, and a direct policy space update mechanism.

[0127] The THPPO proposed in this invention extends the standard PPO to support hybrid action space and dynamic state representation, thereby achieving robust policy updates in dynamically changing environments. The algorithm captures the temporal dependencies in variable-length state sequences by integrating LSTM modules, and adopts a parallel policy network architecture to jointly optimize discrete and continuous actions by directly updating in the policy space, further improving the stability of hybrid action learning. Experimental results show that compared with other deep reinforcement learning algorithms and optimization methods such as PSO and TS, the THPPO algorithm shows significant advantages in learning efficiency, stability, dynamic adaptability and generalization ability. The proposed method breaks through the limitations of the traditional static optimization paradigm by establishing a dynamic optimization mechanism that couples environmental perception with a closed loop of policy generation, and can achieve real-time adaptation to adversarial situations.

[0128] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in the present invention, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0129] In one embodiment, Figure 6 As shown, a device for optimizing an intermittent non-uniform sampling and forwarding interference strategy is provided, comprising:

[0130] A target construction module is used to construct the optimization target of the intermittent non-uniform sampling and forwarding jamming countermeasure process by maximizing the mean absolute deviation of the reference unit power, sampling pulse width sequence, and forwarding number sequence in the jamming round;

[0131] A model building module is used to model the intermittent non-uniform sampling and forwarding jamming countermeasure process as a parameterized action Markov decision process; each time step of the parameterized action Markov decision process includes the current state, action, state transition probability and reward value; the state includes the radar pulse width and the action of the historical time step; the action includes the sampling pulse width of the jamming sub-pulse at the current time step and the corresponding number of forwarding; the state transition probability includes the probability of the jammer transitioning from the current state to the next state after taking the jamming action; the reward value includes the target reward at the end of the jamming round or the potential reward before the end of the jamming round; the potential reward includes the change in the jamming signal distribution, the pulse width difference between the current sampling pulse width and the historical sampling pulse width, and the forwarding number difference between the current forwarding number and the historical forwarding number; the target reward includes the minimum average power of the reference unit, the average absolute deviation of the sampling pulse width and the forwarding number;

[0132] The target reconstruction module is used to reconstruct the optimization target according to the parameterized action Markov decision process to maximize the expected cumulative reward of the interference decision trajectory obtained by sampling the interference strategy to obtain the reconstructed optimization target; the interference decision trajectory includes a state sequence, an action sequence, and a reward sequence;

[0133] The target solving module is used to solve the reconstruction optimization target and obtain the optimized interference strategy.

[0134] In one embodiment, it is also used to construct an interference decision model for predicting the action at each time step in each interference round; the interference decision model includes a two-branch actor and a global critic, the two-branch actor includes a discrete sub-actor and a continuous sub-actor, wherein the input layer of each network includes an LSTM layer, the global critic network is used to calculate the state value under the current state, and the advantage estimate calculated by the state value is used to guide the strategy optimization of the discrete sub-actor network and the continuous sub-actor network, the discrete sub-actor network and the continuous sub-actor network are used to select the number of forwardings and the sampling pulse width under the input state from the discrete action space and the continuous action space respectively; the number of forwardings is obtained by sampling according to the probability distribution of the number of forwardings, and the sampling pulse width is obtained by matching the sampling pulse width Gaussian distribution parameters corresponding to the number of forwardings; the interference decision model is trained according to a preset loss function to obtain a trained two-branch actor, and the real-time state is input into the trained two-branch actor to obtain an optimized interference decision.

[0135] In one embodiment, the global critic includes an LSTM layer, multiple fully connected layers, and an output layer.

[0136] In one embodiment, the discrete sub-actor includes an LSTM layer, multiple fully connected layers, a classification layer, and an output layer, and the continuous sub-actor includes an LSTM layer, multiple fully connected layers, and an output layer.

[0137] In one embodiment, it is further used to perform time width calibration on the sampling pulse width of the last interference sub-pulse in each interference round.

[0138] In one embodiment, when the number of completed interference rounds reaches the model training frequency, the interference decision model is updated based on the sampled experience tuple and a preset loss function; the experience tuple includes the state, state transition probability, reward value, and the number of forwarding times and sampling pulse width output by the interference decision model; the updated interference decision model is used to execute the next interference round until the preset number of training times is met, thereby obtaining a trained dual-branch actor.

[0139] In one embodiment, it is also used to calculate the discounted return corresponding to the state at each time step at the end of the interference round; the discounted return is used to calculate the loss function of the global critic during model training; a generalized advantage estimate is performed based on the state value, reward and state value of each time step and the next time step to obtain an advantage estimate; the advantage estimate is used to calculate the objective function of the discrete sub-actor and the continuous sub-actor during model training.

[0140] In one embodiment, it is also used to maximize the loss function of the global critic based on the sampled experience tuples; the loss function of the global critic is obtained based on the mean square error of the state value and the discounted return; the objective functions of the discrete sub-actor and the continuous sub-actor are minimized based on the sampled experience tuples to update the interference decision model; the objective function of the discrete sub-actor is obtained based on the PPO clipping product of the probability ratio of the new and old strategies in the discrete action space and the advantage estimate, and the objective function of the continuous sub-actor is obtained based on the PPO clipping product of the probability density ratio of the new and old strategies in the continuous action space and the advantage estimate.

[0141] Regarding the specific definition of the intermittent non-uniform sampling and forwarding interference strategy optimization device, please refer to the definition of the intermittent non-uniform sampling and forwarding interference strategy optimization method above, and will not be repeated here. The various modules in the above-mentioned intermittent non-uniform sampling and forwarding interference strategy optimization device can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0142] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, an intermittent non-uniform sampling forwarding interference strategy optimization method is implemented. The display screen of the computer device can be a liquid crystal display or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a key, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.

[0143] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0144] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the method in the above embodiment when executing the computer program.

[0145] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0146] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and such modifications and improvements are intended to fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for optimizing an intermittent non-uniform sampling and forwarding interference strategy, characterized in that: The method comprises: The optimization objective of the intermittent non-uniform sampling and forwarding jamming countermeasure process is constructed by maximizing the mean absolute deviation of the reference unit power, sampling pulse width sequence, and forwarding number sequence in the jamming round. The intermittent non-uniform sampling and forwarding interference countermeasure process is modeled as a parameterized action Markov decision process; each time step of the parameterized action Markov decision process includes the current state, action, state transition probability and reward value; the state includes the radar pulse width and the action of the historical time step; the action includes the sampling pulse width of the interference sub-pulse at the current time step and the corresponding number of forwarding; the state transition probability includes the probability of the jammer transferring from the current state to the next state after taking the interference action; the reward value includes the target reward at the end of the interference round or the potential reward before the end of the interference round; the potential reward includes the change in the interference signal distribution, the pulse width difference between the current sampling pulse width and the historical sampling pulse width, and the forwarding number difference between the current forwarding number and the historical forwarding number; the target reward includes the minimum average power of the reference unit, the average absolute deviation of the sampling pulse width and the forwarding number; Reconstructing the optimization objective according to the parameterized action Markov decision process to maximize the expected cumulative reward of the interference decision trajectory obtained by sampling the interference strategy to obtain a reconstructed optimization objective; the interference decision trajectory includes a state sequence, an action sequence, and a reward sequence; Solve the reconstruction optimization objective to obtain an optimized interference strategy.

2. The method according to claim 1, characterized in that Solving the reconstruction optimization objective to obtain the optimized interference strategy includes: Construct an interference decision model for predicting the action at each time step in each interference round; the interference decision model includes a dual-branch actor and a global critic, the dual-branch actor includes a discrete sub-actor and a continuous sub-actor, wherein the input layer of each network includes an LSTM layer, the global critic network is used to calculate the state value under the current state, and the advantage estimate calculated by the state value is used to guide the strategy optimization of the discrete sub-actor network and the continuous sub-actor network, the discrete sub-actor network and the continuous sub-actor network are used to select the number of forwarding times and the sampling pulse width under the input state from the discrete action space and the continuous action space, respectively; the number of forwarding times is obtained by sampling according to the probability distribution of the number of forwarding times, and the sampling pulse width is obtained by matching the sampling pulse width Gaussian distribution parameters corresponding to the number of forwarding times; The interference decision model is trained according to a preset loss function to obtain a trained dual-branch actor, and the real-time state is input into the trained dual-branch actor to obtain an optimized interference decision.

3. The method according to claim 2, characterized in that The global critic includes an LSTM layer, multiple fully connected layers and an output layer.

4. The method according to claim 2, characterized in that The discrete sub-actor includes an LSTM layer, multiple fully connected layers, a classification layer and an output layer, and the continuous sub-actor includes an LSTM layer, multiple fully connected layers and an output layer.

5. The method according to claim 2, characterized in that After assisting the jammer in making an interference decision according to the trained interference decision model, the method further includes: The sampling pulse width of the last interference sub-pulse in each interference round is calibrated.

6. The method according to claim 2, characterized in that The interference decision model is trained according to a preset loss function to obtain a trained dual-branch actor including: When the number of completed interference rounds reaches the model training frequency, the interference decision model is updated based on the sampled experience tuple and the pre-set loss function; the experience tuple includes the state, state transition probability, reward value, and the number of forwarding times and sampling pulse width output by the interference decision model; The updated interference decision model is used to execute the next interference round until the preset number of training times is met, and a trained two-branch actor is obtained.

7. The method according to claim 6, characterized in that The method further comprises: At the end of the interference round, the discounted return corresponding to the state at each time step is calculated; the discounted return is used to calculate the loss function of the global critic during model training; A generalized advantage estimate is performed based on the state value, reward, and state value of each time step to obtain an advantage estimate; the advantage estimate is used to calculate the objective function of the discrete sub-actor and the continuous sub-actor during model training.

8. The method according to claim 6, characterized in that The updating of the interference decision model according to the sampled experience tuple and the preset loss function includes: Maximizing a global critic loss function based on the sampled experience tuples; the global critic loss function is obtained based on the mean square error of the state value and the discounted return; The objective functions of the discrete sub-actor and the continuous sub-actor are minimized according to the sampled experience tuple to update the interference decision model; the objective function of the discrete sub-actor is obtained according to the probability ratio of the new and old strategies in the discrete action space and the PPO clipping product of the advantage estimate, and the objective function of the continuous sub-actor is obtained according to the probability density ratio of the new and old strategies in the continuous action space and the PPO clipping product of the advantage estimate.

9. An intermittent non-uniform sampling and forwarding interference strategy optimization device, characterized in that: The device comprises: A target construction module is used to construct the optimization target of the intermittent non-uniform sampling and forwarding jamming countermeasure process by maximizing the mean absolute deviation of the reference unit power, sampling pulse width sequence, and forwarding number sequence in the jamming round; A model building module is used to model the intermittent non-uniform sampling and forwarding interference countermeasure process as a parameterized action Markov decision process; each time step of the parameterized action Markov decision process includes the current state, action, state transition probability and reward value; the state includes the radar pulse width and the action of the historical time step; the action includes the sampling pulse width of the interference sub-pulse at the current time step and the corresponding number of forwarding; the state transition probability includes the probability of the jammer transferring from the current state to the next state after taking the interference action; the reward value includes the target reward at the end of the interference round or the potential reward before the end of the interference round; the potential reward includes the change in the interference signal distribution, the pulse width difference between the current sampling pulse width and the historical sampling pulse width, and the forwarding number difference between the current forwarding number and the historical forwarding number; the target reward includes the minimum average power of the reference unit, the average absolute deviation of the sampling pulse width and the forwarding number; A target reconstruction module is used to reconstruct the optimization target according to the parameterized action Markov decision process to maximize the expected cumulative reward of the interference decision trajectory obtained by sampling the interference strategy to obtain a reconstructed optimization target; the interference decision trajectory includes a state sequence, an action sequence and a reward sequence; The target solving module is used to solve the reconstruction optimization target and obtain the optimized interference strategy.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.