Intelligent jamming decision method and system based on prior knowledge embedded LSTM-PPO model

By introducing the LSTM-PPO model with prior knowledge embedding into the multi-function radar jamming decision-making and utilizing the reshaped reward function and the temporal feature extraction capability of LSTM, the problem of traditional methods relying on prior data is solved, and fast convergence and efficient jamming decision-making are achieved.

CN118818440BActive Publication Date: 2025-10-17ZHEJIANG SCI-TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410785431.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-18
Publication Date
2025-10-17
Estimated Expiration
2044-06-18

AI Technical Summary

Technical Problem

Traditional interference decision-making methods rely on prior data and are difficult to achieve real-time and effectiveness in multi-function radar scenarios. Reinforcement learning models have problems with low learning efficiency due to large data requirements and full-domain search.

Method used

The LSTM-PPO model based on prior knowledge embedding is adopted to guide the agent learning by reshaping the reward function. Combined with the temporal feature extraction capability of LSTM, the accuracy and stability of interference decision-making are improved.

Benefits of technology

The algorithm convergence speed and stability of multi-function radar jamming decision-making are significantly improved, the effectiveness and efficiency of jamming decision-making are enhanced, and the optimal jamming strategy can be quickly obtained in complex electromagnetic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118818440B_ABST
    Figure CN118818440B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent interference decision-making method and system based on an LSTM-PPO model embedded with prior knowledge, and belongs to the fields of artificial intelligence and machine learning; the method comprises the following steps: modeling a multifunctional radar environment (MFR) to obtain an environment model; defining an MFR interference decision-making problem as a Markov decision process; embedding prior knowledge in a PPO model in the form of a reshaped reward based on a potential function of the environment model, so as to guide an agent to converge rapidly; and using an LSTM agent PPO algorithm to embed a reinforcement learning model, so as to capture dynamic characteristics of echo data, effectively describe radar working states, and improve interference decision-making precision and stability. The application has high decision-making efficiency and effectiveness, and can efficiently and stably achieve a multifunctional radar interference strategy. The application has significant advantages in terms of convergence speed, stability and performance of executing an interference decision-making under a multifunctional radar environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence and machine learning, and particularly relates to an intelligent interference decision-making method and system based on a LSTM-PPO model embedded with prior knowledge. BACKGROUND

[0002] Multifunctional radar (MFR) can perform search, tracking, identification and other functions in modern battlefield based on its flexible waveform transformation, agile beam scanning and strong electromagnetic countermeasure characteristics, and plays a crucial role in reconnaissance, attack and defense and other combat tasks, and is therefore regarded as the core and key equipment of the electromagnetic spectrum equipment system. Therefore, effectively interfering with MFR to weaken its combat effectiveness has always been a research hotspot in the field of military equipment. The interference process mainly includes electromagnetic perception, interference decision-making and interference effect evaluation, and interference decision-making is the key link of the process, which aims to efficiently formulate interference strategies based on the data obtained in the reconnaissance stage to accurately configure interference patterns and corresponding interference resources, and then effectively execute interference tasks, and finally significantly reduce the threat of enemy radar.

[0003] Traditional interference decision-making methods mainly rely on "template matching", "game decision-making" and "reasoning decision-making" strategies, which all need to rely on prior data such as radar coding mode, power allocation, etc. However, MFR has strong adaptive ability, and the transmitted waveform is flexible and variable. Therefore, the traditional decision-making method based on prior data analysis to obtain interference strategies in the MFR scene faces the problem of difficulty in obtaining prior data, which significantly reduces the real-time and effectiveness of interference decision-making. Therefore, it is particularly important to develop an interference decision-making method that does not rely too much on prior data.

[0004] To solve this problem, reinforcement learning (RL) technology based on "trial and error" to learn the best strategy under the condition of lack of prior data is applied to the field of MFR interference decision-making. Reinforcement learning-based interference decision-making can use rewards as feedback to effectively solve the interference evaluation problem, and the "trial and error" method can significantly reduce the dependence of decision-making on prior data. The autonomous adaptive learning of the agent can effectively improve the intelligent level and efficiency of decision-making, thereby providing a new solution for MFR interference decision-making in complex electromagnetic scenarios.

[0005] MFR interference decision-making model based on deep reinforcement learning has developed rapidly in recent years, but there are still key problems such as decision effectiveness and efficiency that need to be solved. Since the reinforcement learning model is usually based on a deep learning network, it has typical characteristics of data-driven models such as large data sample requirement and global search, which leads to low learning efficiency. Therefore, in order to improve the decision efficiency of the reinforcement learning model, the existing prior knowledge must be combined to guide the model optimization, so that the model quickly converges to the optimal strategy. SUMMARY

[0006] The present application aims to provide an intelligent interference decision-making method and system based on LSTM-PPO model embedded with prior knowledge, which uses LSTM to embed a deep reinforcement learning model to effectively extract the essential features of time series data and improve the effectiveness of interference decision-making algorithms based on reinforcement learning, and embeds interference-related prior knowledge into the reinforcement learning model to guide its rapid convergence and improve the efficiency of interference decision-making models based on reinforcement learning.

[0007] According to a first aspect of the embodiments of the present disclosure, an intelligent interference decision-making method based on LSTM-PPO model embedded with prior knowledge is provided, comprising the following steps:

[0008] Modeling the multi-function radar environment MFR to obtain an environment model;

[0009] Defining the MFR interference decision-making problem as a Markov Decision Process (MDP);

[0010] The reshaped reward theory based on the potential function of the environment model embeds prior knowledge in the PPO model in the form of reshaped reward to guide the agent to quickly converge;

[0011] Using LSTM agent PPO algorithm to embed the reinforcement learning model is used to capture the dynamic characteristics of echo data to effectively describe the radar working state and improve the interference decision-making precision and stability.

[0012] According to a second aspect of the embodiments of the present disclosure, an intelligent interference decision-making system based on LSTM-PPO model embedded with prior knowledge is provided, comprising:

[0013] The modeling module models the multi-function radar environment MFR to obtain an environment model;

[0014] The definition module defines the MFR interference decision-making problem as a Markov Decision Process (MDP);

[0015] The embedding module embeds prior knowledge in the form of reshaped rewards into the PPO model based on a potential function of an environmental model, so as to guide the agent to quickly converge.

[0016] The capturing module embeds a reinforcement learning model using an LSTM agent PPO algorithm, and is used for capturing dynamic characteristics of echo data to effectively depict a radar working state, and improves interference decision precision and stability.

[0017] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, which comprises a memory, a processor and a computer program stored in the memory and running on the memory, and the processor implements the intelligent interference decision method of the LSTM-PPO model embedded based on prior knowledge when executing the program.

[0018] According to a fourth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, which stores a computer program, and the program is executed by a processor to implement the intelligent interference decision method of the LSTM-PPO model embedded based on prior knowledge.

[0019] Compared with the prior art, the method has the following advantages: firstly, the interference-related prior information is embedded into the reward function by using the reward reshaping theory based on the potential function, so as to more effectively guide the policy learning of the agent. Then, based on the excellent time sequence feature extraction capability of the LSTM, the accurate dynamic characteristics of the radar sequence data are captured to effectively depict the radar working state. Finally, the extracted dynamic characteristics are input into the reinforcement learning model based on the policy gradient, and the effective interference strategy can be quickly obtained through the guidance of the embedded prior knowledge.

[0020] The present application has significant advantages in convergence speed, stability and performance of executing interference decision under the environment of a multifunctional radar. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0022] Figure 1 LSTM embedding PPO model diagram of the present application;

[0023] Figure 2 Method overall framework flowchart of the present application;

[0024] Figure 3Fig. 4 is a comparison chart of average rewards obtained by interference decision algorithms based on different reinforcement learning models;

[0025] Figure 4 Fig. 5 is a comparison chart of rewards obtained by PPO interference decision models constructed based on different deep networks;

[0026] Figure 5 Fig. 6 is a comparison chart of rewards obtained by prior knowledge embedding. DETAILED DESCRIPTION

[0027] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.

[0028] It should be noted that the following detailed description is illustrative only and is intended to provide further description of the present application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0029] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments according to the present application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, elements, components, and / or groups thereof, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.

[0030] It should be noted that the flowchart and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of methods and systems according to various embodiments of the present disclosure. It should also be noted that each block in the flowchart or block diagrams can represent a module, a segment, or a portion of code, which can include one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession can in fact be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the flowchart or block diagrams and combinations of blocks in the flowchart or block diagrams can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.

[0031] Embodiment One:

[0032] The embodiment provides an intelligent interference decision method based on a prior knowledge embedded LSTM-PPO model, including the following steps:

[0033] Step 1: Model the multi-function radar environment MFR to obtain an environment model;

[0034] Specifically, the MFR state, complex electromagnetic environment and jamming action are described as a mathematical model containing the following elements: MFR finite working state set S (s∈S), jammer action set A (a∈A), reward set R that depends on state transition, and benefit function R(s t |a t ,s t+1 ), the state transition probability P(s t+1 |s t ,a t ) characterizes the environment model. The jammer's jamming pattern forces the radar state to shift, thereby obtaining corresponding benefits for the jamming decision system. Iterative attempts aim to maximize the cumulative expected reward, thereby obtaining the optimal jamming strategy and ultimately achieving the jamming goal.

[0035] Step 2: Define the MFR interference decision problem as a Markov decision process MDP:

[0036] Specifically, the Markov decision process MDP can be defined by the four-tuple {S, A, P, R}. Among them, S is the state set, and the radar working state is recorded as s i ω ,i=0,1,2,...,N s , threat level from 0 to N s Descending successively, N s is the number of radar working states, ω is the radar waveform unit; A is the action set, which can be expressed as a t ={jam0,jam2,...jam I}, I is the number of actions; P is the number of actions from the current state s t Take action a t Go to the next state s t+1 The transition probability P(s t+1 |s t ,a t ); R is the state s t 、Action a t and the next state s t+1 The set of benefits obtained under given conditions is specifically represented by R1 and R2:

[0037]

[0038] Among them, Δ(s t+1 -s t ) is the change in threat level after the radar is affected by interference and migrates to a new radar state. +100 means that the radar state has migrated to the optimal target state. s endR2 is the threat degree change of state parameters affected by jamming and before jamming, if the threat degree increases or remains unchanged, then the penalty is 1, otherwise the reward is 1, represents the norm of the state parameter normalization result.

[0039] Step 3: The reshaped reward theory based on the potential function of the environment model embeds prior knowledge in the form of reshaped reward into the PPO model to guide the agent to quickly converge:

[0040] Specifically, according to the Bellman equation, the optimal state-value function can be expressed as follows:

[0041]

[0042] In the formula: is the optimal value function, is the expectation of the action taken in the next radar state, s' is the next radar state, γ is the discount factor, and a' is the next jammer action.

[0043] Based on the definition of the potential function, we have:

[0044]

[0045] In the formula: as defined above; φ(s) is the potential function; E s' as defined above, this is a shorthand; γφ(s') is the potential function of the next radar state multiplied by the discount factor.

[0046] Based on the difference form of the potential function F(s, a, s') = γφ(s') - φ(s), we have:

[0047]

[0048] Thus, when M' reaches the optimal strategy, the action-value function of M' is satisfies the following conditions:

[0049]

[0050] In the formula: represents the optimal strategy of M'.

[0051] It is proved that the optimal strategy of M' is the same as M, which shows that the potential function is only related to the state and has no effect on the action selection under the same state, so the reshaped reward theory based on the potential function does not change the optimal strategy of reinforcement learning.

[0052] The potential function based on the environment model can be obtained by solving the inverse Bellman equation. Specifically, the state s tThe value function can be expressed as follows:

[0053]

[0054] Where: V π (s t+1 ) is the state s obtained under the strategy π t+1 value function, and γ is the discount factor.

[0055] Quantized state s t Relative to the target state s end The potential energy difference of the reaction state transfer energy change is designed as follows Φ(s t ):

[0056]

[0057] Where: U(s t+1 ) is the radar state s t+1 The reward reshaping function under .

[0058] The above formula reverses the reward function R t To ensure the reshaping of the reward function U(s t ) decreases monotonically as the agent approaches the target state, and the introduction of the discount factor γ makes the potential energy function decay reasonably over time.

[0059] Based on the above, we can know that when the agent approaches the target state, the potential energy function value gradually decreases, and the constructed potential energy function Φ(s t ) is non-negative, so in the reshaping reward function, when the agent state conforms to the domain prior information, the reshaping reward function takes the potential energy lost when the agent approaches the target state as a positive reward, and the potential energy gained when it moves away from the prior information target state as a negative reward, thereby ensuring that the agent can learn the strategy along the direction of the prior information. The following reshaping reward function can be constructed:

[0060]

[0061] Combined with the above constructed reward functions R1, R2 and the reshaped reward function U(s t ,s t+1 ) to obtain the following new reward function to accelerate the learning process of the agent:

[0062] R′ 1,2 (s t ,s t+1 )=R 1,2 (s t ,s t+1 )+α·ΔU(s t ,s t+1 ) (10)

[0063] Where R 1,2 (s t ,s t+1 ) is the original reward function, ΔU(s t ,s t+1 ) is the potential energy change, and α is a scalar weight used to adjust the impact of the potential energy change on the reshaping reward.

[0064] Step 4: Use the LSTM proxy PPO algorithm to embed a reinforcement learning model to capture the dynamic characteristics of the echo data to effectively characterize the radar operating status and improve the accuracy and stability of interference decision-making:

[0065] Specifically, the reconnaissance sequence data is first preprocessed based on a fully connected network, and the network weights are orthogonally initialized via equations (11) and (12). This initialization strategy aims to improve the stability of network training and avoid the problem of gradient disappearance or explosion.

[0066] z t =W0s t +b0 (11)

[0067] Among them, W0 is the weight matrix, which can be orthogonally initialized as W0=Q0D0, Q0 is an orthogonal matrix, D0 is a diagonal matrix, b0 is a bias vector, s t is the state vector.

[0068] Then, z is activated based on the nonlinear function σ t , we can get

[0069] h t =σ(z t )=σ(W0s t +b0) (12)

[0070] Activated h t It can be used as LSTM input, and the following hidden state can be obtained after LSTM processing:

[0071] h′ t =lstm(h′ t-1 ,h t ) (13)

[0072] Where h′ t is the hidden state of LSTM at time step t, h′ t-1 is the hidden state of the previous time step. The hidden state of LSTM allows the two core networks of the PPO framework to be constructed, namely: the value network (Critic Network) and the policy network (Actor Network)

[0073] The resulting hidden state is then input into the value network and the policy network to obtain:

[0074]

[0075] where V(s t ) is the value network's resulting state s t long-term value, and b v are the weight vector and bias of the value network respectively, π(a t |s t ) is the policy network's resulting policy, and b a are the weight vector and bias of the policy network respectively, and Softmax is the classification function.

[0076] Finally, the policy network updates its network through its loss function L CLIP (θ) and based on the gradient descent algorithm, and the value network updates its network through the loss function L GAE (θ v ) and based on the gradient descent algorithm, and the specific loss function can be expressed as:

[0077]

[0078] where π θ (a|s) is the probability of selecting action a in state s, is the advantage function, ε is the clipping coefficient for limiting the scale of policy updates, is the advantage function estimate value, and θ is the policy network parameter.

[0079]

[0080] where θ v is the value network parameter, is the estimated advantage value output by the value network, N is the number of samples, GAE is the generalized advantage estimation function, γ is the discount factor, and λ is the weighting parameter for controlling bias and variance.

[0081] Embodiment Two:

[0082] The embodiment provides an intelligent interference decision system based on an LSTM-PPO model embedded with prior knowledge, which comprises:

[0083] A modeling module models a multi-function radar environment (MFR) to obtain an environment model;

[0084] A definition module defines the MFR interference decision problem as a Markov decision process (MDP);

[0085] The embedding module embeds prior knowledge in the form of reshaped rewards into the PPO model based on a reshaped reward theory of a potential function of an environment model, to guide the agent to quickly converge;

[0086] The capturing module embeds a reinforcement learning model using an LSTM agent PPO algorithm, which is used to capture the dynamic characteristics of echo data to effectively depict the radar working state, and improve the interference decision precision and stability.

[0087] Embodiment three:

[0088] An electronic device comprises a memory, a processor and a computer program stored on the memory, and the processor executes the program to implement the above-mentioned intelligent interference decision method of the LSTM-PPO model embedded based on prior knowledge, comprising:

[0089] Modeling the multi-function radar environment (MFR) to obtain an environment model;

[0090] Defining the MFR interference decision problem as a Markov decision process (MDP);

[0091] The reshaped reward theory of the potential function of the environment model embeds prior knowledge in the form of reshaped rewards into the PPO model, to guide the agent to quickly converge;

[0092] The LSTM agent PPO algorithm is used to embed a reinforcement learning model, which is used to capture the dynamic characteristics of echo data to effectively depict the radar working state, and improve the interference decision precision and stability.

[0093] Embodiment four:

[0094] A computer readable storage medium has a computer program stored thereon, and the program is executed by a processor to implement the above-mentioned intelligent interference decision method of the LSTM-PPO model embedded based on prior knowledge, comprising:

[0095] Modeling the multi-function radar environment (MFR) to obtain an environment model;

[0096] Defining the MFR interference decision problem as a Markov decision process (MDP);

[0097] The reshaped reward theory of the potential function of the environment model embeds prior knowledge in the form of reshaped rewards into the PPO model, to guide the agent to quickly converge;

[0098] The LSTM agent PPO algorithm is used to embed a reinforcement learning model, which is used to capture the dynamic characteristics of echo data to effectively depict the radar working state, and improve the interference decision precision and stability.

[0099] The effects of the present application can be further illustrated by the following simulation:

[0100] The simulation conditions are as follows: the learning rate η = 0.01, the reward discount factor γ = 0.99, the decay weight β = 0.9, the GAE advantage function weight λ = 0.95, the initial exploration rate ε = 0.1, the terminal exploration rate ε = 0.9, the PPO clipping coefficient δ = 0.2, the PPO entropy regularization term H = 0.01, the balance coefficient α of priori information = 32, the total amount of experience replay pool D is 1000, and the network parameters are updated every 15 steps or at the end of the current network round. The parameter settings of the comparative reinforcement learning algorithm are as follows: the TRPO conjugate gradient algorithm optimization iteration number is 15, the damping term is 0.1, the control policy update amplitude value is 0.01, and the remaining parameters are the same as those of PPO; the DQN algorithm learning rate η = 0.01, the reward discount factor γ = 0.99, and the total amount of experience replay pool D is 1000; the network parameters are updated every 15 steps; and the learning rate η = 0.01 and the reward discount factor γ = 0.99 in the Q-Learning algorithm. In the experiment, the end condition of each iteration is that the radar state enters the terminal state or the number of trial and error reaches 300.

[0101] Simulation content:

[0102] Simulation 1: Comparison of rewards obtained by different reinforcement learning algorithms. The average rewards obtained by interference decision algorithms based on different reinforcement learning models are compared, as shown in Figure 3 The rewards obtained by PPO interference decision models based on different deep networks are compared, as shown in Figure 4 The rewards obtained by embedding priori knowledge are compared, as shown in Figure 5

[0103] As can be seen from Figure 3 , the average reward values obtained by the above four algorithms increase with the number of rounds, indicating that the effectiveness of the obtained strategy increases with the number of trial and error. Secondly, when Q-Learning and DQN algorithms balance exploration and utilization, the intelligent agent will appear obvious fluctuations in rewards in the short term. This volatility is an inevitable phenomenon in the process of the intelligent agent constantly trying new strategies to find better solutions. In addition, with the increase of the number of rounds, the DQN algorithm may fall into a local optimal solution, which not only causes the obtained reward to fluctuate continuously, but also limits the improvement of the effectiveness of the intelligent agent strategy. Furthermore, in the application of mapping from discrete to continuous space, due to the existence of estimation error and approximation error, the TRPO algorithm cannot always guarantee that the reward obtained by the new strategy is higher than that by the old strategy. Therefore, with the increase of the number of rounds, the reward value of the TRPO algorithm may decrease, which is caused by the above-mentioned errors that cannot be completely avoided in the process of strategy iteration. In addition, compared with the comparative reinforcement learning algorithm, the PPO algorithm not only converges quickly to a higher reward value, but also has smaller fluctuations in the reward obtained with the increase of the number of rounds, indicating that the interference decision algorithm based on PPO has higher decision effectiveness, efficiency and stability.​

[0104] Depend on Figure 4 It can be seen that the average reward value obtained by the PPO interference decision model constructed with CNN as the basic architecture is significantly lower than that of LSTM. This can be attributed to the fact that although CNN has unique advantages in spatial feature extraction, the interference decision model does not involve the spatial characteristics of the data, and therefore its decision-making efficiency and effectiveness in the policy optimization process are relatively limited. In contrast, the PPO interference decision model embedded with LSTM has higher decision-making efficiency and effectiveness. This is because the interference decision task has a sequential nature, and LSTM has a significant advantage in handling temporal dependency-related problems due to its inherent long-term memory mechanism. Its gating structure enables the network to effectively capture and fully utilize historical decision information, thereby improving the efficiency and effectiveness of interference decision-making.

[0105] Depend on Figure 4 、 5 It can be seen that compared to the interference decision-making model lacking prior information, the interference decision-making model embedded with prior information has higher decision-making efficiency and effectiveness, indicating that prior information can guide the decision-making model to quickly converge to a better strategy and obtain higher strategic returns. Furthermore, due to the excellent temporal feature extraction capabilities of LSTM, compared with CNN and traditional linear models, LSTM-PPO demonstrates better decision-making effectiveness in the early stages of strategy optimization. In addition, the LSTM-PPO with embedded prior knowledge (i.e., PK-LSTM-PPO) consistently achieves higher average reward values ​​than PK-CNN-PPO and PK-PPO in the middle and late stages of strategy optimization. This shows that compared with the comparison algorithms, the proposed LSTM-PPO algorithm with embedded prior knowledge can converge to a better and more stable interference strategy more quickly.

[0106] Simulation 2: Convergence performance comparison simulation, as shown in Table 1.

[0107] To obtain the effective comparison of the convergence performance of the jamming decision algorithm based on different deep reinforcement models, the number of training rounds is increased from 500 to 1000 to more comprehensively evaluate the convergence performance of different algorithms over a long time span. Other factors, including the configuration of the computing hardware, the network architecture, and the network size, are consistent with the above experiment. Thus, the average time used in each round is shown in Table 3, where the algorithm can be divided into two parts before and after embedding prior knowledge. As shown in Table 3, without embedding prior knowledge in the algorithm, LSTM-PPO has a faster convergence speed, and its convergence time is improved by 38.1% relative to the slower TRPO algorithm. Furthermore, compared with the model without embedding prior knowledge, the jamming decision efficiency of the algorithm with embedding prior knowledge is significantly improved. In addition, compared with other comparative algorithms, the PK-LSTM-PPO algorithm has better convergence performance, and its convergence time is improved by 7.64 times relative to the LSTM-PPO algorithm without embedding prior knowledge, and is improved by 34.6% and 47% relative to the PK-PPO and PK-CNN-PPO algorithms, respectively. Thus, based on the excellent time series feature extraction capability of LSTM and the strong convergence guiding capability of the domain prior knowledge, the proposed algorithm has better convergence performance compared with the comparative algorithms, and thus can more effectively counteract the complex and variable electromagnetic environment.

[0108] Table 1 Time used for convergence in each round of different reinforcement learning algorithms

[0109]

[0110] Simulation 3: Decision path comparison simulation, as shown in Table 2.

[0111] Based on the prior information, the following two shortest jamming paths can be obtained:

[0112]

[0113] wherein, is the radar state.

[0114] Based on the principle of interference equipment being free from dangerous state, the strategy that makes the radar state transfer to the state with the lowest threat as soon as possible along the direction of decreasing threat level is defined as the correct interference strategy. Table 2 lists the correct interference strategies of the shortest interference paths obtained by each algorithm after 500 rounds, the interference strategies that do not obtain the shortest paths, and the failure rate of the interference strategies. As can be seen from Table 2, compared with various variant algorithms of PPO, the traditional reinforcement learning algorithms such as Q-Learning, DQN, TRPO and PPO have a higher failure rate of interference, indicating that the robustness of the traditional algorithms in selecting interference paths is lower. Furthermore, the interference failure rate of the PPO algorithm based on LSTM and CNN is lower, indicating that the deep network model can extract effective features of echo data, thereby significantly improving the decision efficiency and effectiveness. In addition, the embedding of prior knowledge can guide the PPO and its variant algorithms to the shortest path, making them tend to the shortest and effective interference path, thereby shortening the interference path and reducing the probability of interference failure.

[0115] Table 2 Comparison of decision paths

[0116]

[0117] Those skilled in the art should understand that each module or each step of the present disclosure described above can be realized by a general computer device, and alternatively, they can be realized by program codes executable by a computing device, so that they can be stored in a storage device and executed by a computing device, or they can be respectively manufactured into each integrated circuit module, or a plurality of modules or steps thereof can be manufactured into a single integrated circuit module. The present disclosure is not limited to any specific combination of hardware and software.

[0118] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

[0119] The above describes the specific embodiments of the present disclosure in conjunction with the accompanying drawings, but is not intended to limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications or changes made on the basis of the technical solutions of the present disclosure without inventive labor are still within the protection scope of the present disclosure.

Claims

1. An intelligent interference decision-making method based on the LSTM-PPO model embedded with prior knowledge, characterized by: Specifically include: Modeling the multi-function radar environment MFR to obtain an environment model; The MFR interference decision problem is defined as a Markov decision process MDP; The reshaping reward theory of the potential energy function based on the environment model embeds prior knowledge into the PPO model in the form of reshaping reward to guide the rapid convergence of the intelligent agent. Specifically, when the intelligent agent approaches the target state, the value of the potential energy function gradually decreases, and the constructed potential energy function Φ(s t ) is non-negative, so in the reshaping reward function, when the agent state conforms to the domain prior information, the reshaping reward function takes the potential energy lost when the agent approaches the target state as a positive reward, and the potential energy gained when it moves away from the prior information target state as a negative reward, thereby ensuring that the agent can learn the strategy along the direction of the prior information. The following reshaping reward function is constructed: Combined with the constructed reward function R1, R2 and the reshaped reward function U(s t ,s t+1 ) to obtain the following new reward function to accelerate the learning process of the agent: R′ 1,2 (s t ,s t+1 )=R 1,2 (s t ,s t+1 )+α ΔU(s t ,s t+1 ) (10) Where R 1,2 (s t ,s t+1 ) is the original reward function, ΔU(s t ,s t+1 ) is the potential energy change, α is the scalar weight used to adjust the impact of potential energy change on the reshaping reward; s t is the current state, s t+1 For the next state; The LSTM proxy PPO algorithm is used to embed a reinforcement learning model to capture the dynamic characteristics of the echo data to effectively characterize the radar working state and improve the accuracy and stability of interference decision-making.

2. The intelligent interference decision-making method based on the LSTM-PPO model embedded with prior knowledge according to claim 1 is characterized in that: The multi-function radar environment MFR is modeled as follows: the MFR state, complex electromagnetic environment and jamming action are described as a mathematical model containing the following elements: MFR finite working state set S (s∈S), jammer action set A (a∈A), reward set R that depends on state transition, and benefit function R(s t |a t ,s t+1 ), the state transition probability P(s t+1 |s t ,a t ) characterizes the environmental model; the jammer launches interference patterns to force the radar state to shift, thereby interfering with the decision-making system to obtain corresponding benefits, and iterative attempts are made to maximize the cumulative expected reward to obtain the optimal jamming strategy and ultimately achieve the jamming goal.

3. The intelligent interference decision-making method based on the LSTM-PPO model embedded with prior knowledge according to claim 1 or 2 is characterized in that: The Markov decision process MDP is defined by the four-tuple {S, A, P, R}, where S is the state set and the radar working state is recorded as s i ω ,i=0,1,2,...,N s , threat level from 0 to N s Descending successively, N s is the number of radar working states, ω is the radar waveform unit; A is the action set, expressed as a t ={jam0,jam2,…jam I }, I is the number of actions; P is the number of actions from the current state s t Take action a t Go to the next state s t+1 The transition probability P(s t+1 |s t ,a t ); R is the state s t 、Action a t and the next state s t+1 The set of benefits obtained under given conditions is specifically represented by R1 and R2: Among them, Δ(s t+1 -s t ) is the change in threat level after the radar is affected by interference and migrates to a new radar state. +100 means that the radar state has migrated to the optimal target state. s end is the target radar state with the lowest threat level, R2 is the threat level change between the state parameter affected by the interference and the parameter before the interference. If the threat level increases or remains unchanged, a penalty of 1 is given, otherwise a reward of 1 is given. represents the norm obtained by normalizing the state parameters.

4. The intelligent interference decision-making method based on the LSTM-PPO model embedded with prior knowledge according to claim 2 is characterized in that: The feasibility verification of the potential energy function's reshaping reward theory is as follows: According to the Bellman equation, the optimal value function is expressed as: Where: is the optimal value function, is the expectation of the action taken at the next radar state, s' is the next radar state, γ is the discount factor, and a' is the next jammer action; Based on the potential energy function definition: Where: Same as above definition; φ(s) is the potential energy function; E s' Same as above definition, here is abbreviated; γφ(s') is the potential energy function of the next radar state multiplied by the discount factor; Based on the potential energy function differential form F(s,a,s')=γφ(s')-φ(s), we can get: It can be seen that when M' reaches the optimal strategy, the action value function of M' is The following conditions must be met: Where: Denote the optimal strategy of M'; It is proved that the optimal strategy of M' is the same as M, indicating that the potential energy function is only related to the state and has no effect on the action selection under the same state. Therefore, the reshaping reward theory of the potential energy function does not change the optimal strategy of reinforcement learning.

5. The intelligent interference decision-making method based on the LSTM-PPO model embedded with prior knowledge according to claim 4 is characterized in that: The potential energy function based on the environmental model is obtained by solving the inverse Bellman equation: Specifically, the state s t The value function is expressed as follows: Where: V π (s t+1 ) is the state s obtained under the strategy π t+1 Value function, γ is the discount factor; Quantized state s t Relative to the target state s end The potential energy difference of the reaction state transfer energy change is designed as follows Φ(s t ): Where: U(s t+1 ) is the radar state s t+1 The reward reshaping function under π(a t |s t ) is the strategy obtained by the strategy network; The above formula reverses the reward function R t To ensure the reshaping of the reward function U(s t ) decreases monotonically as the agent approaches the target state, and the introduction of the discount factor γ makes the potential energy function decay reasonably over time.

6. The intelligent interference decision-making method based on the LSTM-PPO model embedded with prior knowledge according to claim 2 is characterized in that: The LSTM agent PPO algorithm is used to embed the reinforcement learning model, as follows: First, the sequence data obtained by reconnaissance is preprocessed based on the fully connected network, and the network weights are orthogonally initialized using equations (11) and (12): With t =W0s t +b0 (11) Among them, W0 is the weight matrix, which can be orthogonally initialized as W0=Q0D0, Q0 is an orthogonal matrix, D0 is a diagonal matrix, b0 is a bias vector, s t is the state vector; Then, z is activated based on the nonlinear function σ t ,have to h t =σ(z t )=σ(W0s t +b0) (12) Activation h t As LSTM input, the following hidden state is obtained after LSTM processing: h′ t =lstm(h′ t-1 ,h t ) (13) Where h′ t is the hidden state of LSTM at time step t, h′ t-1 is the hidden state of the previous time step; the hidden state of LSTM allows the two core networks of the PPO framework to be constructed, namely: the value network Critic Network and the policy network ActorNetwork; The resulting hidden state is then input into the value network and the policy network to obtain: Where, V(s t ) is the state s obtained by the value network t Long-term value, and b v are the weight vector and bias of the value network, π(a t |s t ) is the strategy obtained by the strategy network, and b a are the weight vector and bias of the policy network respectively.

7. According to the intelligent interference decision-making method based on the LSTM-PPO model embedded with prior knowledge in claim 6, the policy network uses its loss function L CLIP (θ) and update the network based on the gradient descent algorithm. The value network is updated through the loss function L GAE (θ v ) and update the network based on the gradient descent algorithm. The specific loss function is expressed as: Where, π θ (a|s) is the probability of selecting action a in state s, is the advantage function, ε is the clipping coefficient, which is used to limit the strategy update scale, is the estimated value of the advantage function, θ is the policy network parameter; Where θ v is the value network parameter, is the estimated advantage value output by the value network, and N is the number of samples.

8. An intelligent interference decision-making system based on the LSTM-PPO model embedded with prior knowledge, characterized in that: include: A modeling module models the multi-function radar environment MFR to obtain an environment model; Definition module, defining the MFR interference decision problem as a Markov decision process (MDP); Embedding module, based on the reshaping reward theory of the potential energy function of the environment model, embeds prior knowledge into the PPO model in the form of reshaping reward to guide the agent to converge quickly. Specifically, when the agent approaches the target state, the value of the potential energy function gradually decreases, and the constructed potential energy function Φ(s t ) is non-negative, so in the reshaping reward function, when the agent state conforms to the domain prior information, the reshaping reward function takes the potential energy lost when the agent approaches the target state as a positive reward, and the potential energy gained when it moves away from the prior information target state as a negative reward, thereby ensuring that the agent can learn the strategy along the direction of the prior information. The following reshaping reward function is constructed: Combined with the constructed reward function R1, R2 and the reshaped reward function U(s t ,s t+1 ) to obtain the following new reward function to accelerate the learning process of the agent: R′ 1,2 (s t ,s t+1 )=R 1,2 (s t ,s t+1 )+α·ΔU(s t ,s t+1 ) (10) In the formula, R 1,2 (s t ,s t+1 ) is the original reward function, ΔU(s t ,s t+1 ) is the potential energy change, α is the scalar weight used to adjust the impact of potential energy change on the reshaping reward; s t is the current state, s t+1 For the next state; The capture module uses the LSTM proxy PPO algorithm to embed a reinforcement learning model to capture the dynamic characteristics of the echo data to effectively characterize the radar working status and improve the accuracy and stability of interference decision-making.

Citation Information

Patent Citations

  • Optimal strategy obtaining method and device based on double-layer deep reinforcement learning model

    CN114723065A

  • Radar waveform game system construction method and device based on deep reinforcement learning, computer and storage medium

    CN115993582A