Radar Anti-Jamming Strategy Learning Method Based on Knowledge and Model-Based Reinforcement Learning
By introducing knowledge-based and model-based reinforcement learning methods in radar anti-interference strategy learning, using prior information database and weight coefficient optimization, the problem of low sampling efficiency of radar anti-interference strategy design in the existing technology is solved, and more efficient anti-interference strategy learning and better ability to adapt to complex environments is achieved.
Patent Information
- Application Number
- CN202310160800.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-23
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-02-23
AI Technical Summary
The existing radar anti-interference strategy design method based on deep model-free reinforcement learning has low sampling efficiency, requires a large number of training samples to achieve acceptable performance, and it is difficult to adapt to unknown interference in complex environments.
Using a method based on knowledge and model reinforcement learning, a priori information database is constructed by making the radar confront the first jammer with a known multiple interference strategies, the learning model parameters are updated using the first interactive information and the learned anti-interference strategy, the unknown interference strategy is decomposed into the weighted sum of the known interference strategy, the objective function is constructed and the weight coefficient is optimized to obtain the optimal anti-interference strategy.
It improves the efficiency of radar learning anti-jamming strategies, can build a simulation environment based on fewer interactive samples, generate a large number of cheap samples, significantly reduce sample complexity and improve the anti-jamming performance of radar in complex environments.
Smart Images

Figure CN116401556B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of radar, and particularly relates to a radar anti-jamming strategy learning method based on knowledge and model-based reinforcement learning. Background Art
[0002] In recent years, with the continuous improvement of the software and hardware and the intelligent level of jammers, the electromagnetic environment faced by radars has become increasingly complex. Therefore, if a radar only adopts a fixed anti-jamming strategy, it can only cope with certain specific types of interference, which will seriously reduce the anti-jamming performance of the radar.
[0003] To improve the adaptability and learning ability of radars in complex jamming environments, reinforcement learning (RL) has attracted the attention of many researchers. For a given task, reinforcement learning aims to enable an agent to learn an optimal (or near-optimal) solution by interacting with the environment. Different from supervised learning, the agent is not told the "correct" actions to complete the task, and it can only obtain a scalar reward that evaluates the quality of the current action by interacting with the environment. Therefore, reinforcement learning can enable the agent to learn the optimal strategy for completing the given task by itself through the interaction information.
[0004] Currently, the radar anti-jamming strategy design methods based on reinforcement learning mainly focus on the design of the carrier frequency selection strategy of frequency agile (FA) radars. A key problem existing in the existing work is the low sampling efficiency, that is, a large number of samples are required to enable the intelligent radar to reach an acceptable performance. More specifically, the current work is mainly based on deep model-free reinforcement learning. Therefore, learning an effective anti-jamming strategy requires a large number of training samples, which makes it difficult for the radar to adapt to unknown interference in complex environments during online confrontation. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a radar anti-jamming strategy learning method based on knowledge and model-based reinforcement learning. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0006] An embodiment of the present invention provides a radar anti-jamming strategy learning method based on knowledge and model-based reinforcement learning, including the steps of:
[0007] S1. Making the radar confront a first jammer with known multiple jamming strategies to perform anti-jamming strategy learning, and constructing a prior information library by using the first interaction information and the learned anti-jamming strategy;
[0008] S2. Making the radar select actions according to the current strategy to interact with a second jammer with an unknown jamming strategy to collect real experience, and obtaining second interaction information between the radar and the second jammer;
[0009] S3. Update the parameters of the learning model using the first interaction information and the second interaction information;
[0010] S4. Decompose the unknown interference strategy into a weighted sum of known interference strategies in the prior information library using the weight coefficient, and construct the objective function for radar decision-making;
[0011] S5. Measure the KL distance between the unknown interference strategy and the transition probability caused by the known interference strategies in the prior information library using the updated learning model to evaluate the model approximation loss;
[0012] S6. Evaluate the similarity between the unknown interference strategy and the known interference strategies using the model approximation loss, and calculate the weight coefficient;
[0013] S7. Calculate and update the radar anti-jamming strategy using the weight coefficient and the objective function;
[0014] S8. Loop steps S2 - S7 until the radar performance converges or meets the preset requirements to obtain the optimal radar anti-jamming strategy.
[0015] In an embodiment of the present invention, step S2 includes:
[0016] Cause the radar to interact with the second jammer of the unknown interference strategy by selecting actions according to the current strategy to collect real experience, and obtain the second interaction information between the radar and the second jammer;
[0017] Store the second interaction information in the memory pool:
[0018]
[0019] Wherein, represents the memory pool for storing the second interaction information, represents the state information of the sample, represents the action taken by the intelligent radar in the current state, represents the reward received by the intelligent radar after taking the action in the current state, represents the next state reached by the intelligent radar after taking the action in the current state, N inter represents the number of samples collected, and M represents the number of pulses within a CPI.
[0020] In an embodiment of the present invention, step S3 includes:
[0021] Use the first interaction information and the second interaction information to minimize the objective function of the learning model by means of stochastic gradient descent to update the parameters of the learning model, wherein the objective function of the learning model is:
[0022]
[0023] Among them, φ d represents the network parameters of the learning model, represents each known prior information to generate training samples, expresses the learning model.
[0024] In an embodiment of the present invention, step S4 includes:
[0025] Using the weight coefficient to decompose the unknown interference strategy into a weighted sum of the known interference strategies in the prior information library:
[0026]
[0027] Among them, represents the unknown interference strategy, represents the d-th known interference strategy, d = 1, 2,..., D, λ d ≥0, d = 1, 2,..., D and λ d represents the weight of each known interference strategy;
[0028] Based on the prior information library of the radar, combining the decomposition formula of the unknown interference strategy to construct the objective function of radar decision-making:
[0029]
[0030] Among them, π * represents the optimal radar anti-jamming strategy, π represents the anti-jamming strategy of the radar, represents the state value function of the intelligent radar under the d-th known interference, s represents the state information.
[0031] In an embodiment of the present invention, step S5 includes:
[0032] Using the updated learning model to calculate the transition probability caused by the unknown interference strategy and the transition probability caused by the known interference strategy;
[0033] Using the KL distance between the transition probability caused by the unknown interference strategy and the transition probability caused by the known interference strategy to calculate the model approximation loss.
[0034] In an embodiment of the present invention, the model approximation loss is:
[0035]
[0036] Among them, represents the model approximation loss, Denote the d-th known jamming strategy, where d = 1, 2, ..., D. Denote the unknown jamming strategy The sample distribution determined by the transition probability and the strategy adopted by the radar, D KL Denote the KL distance between the transition probability caused by the unknown jamming strategy and the transition probability caused by the known jamming strategy. P(·|s,a) represents the transition probability caused by the known jamming strategy, P d (δ|s,a) represents the transition probability caused by the known jamming strategy.
[0037] In an embodiment of the present invention, step S6 includes:
[0038] Use the model to approximately evaluate the similarity between the unknown jamming strategy and the known jamming strategy, and calculate the weight coefficient by solving the target optimization problem, where the target optimization problem is:
[0039]
[0040] where λ * Denote the optimal weight, λ represents the weight coefficient, λ d Denote the weight of each known jamming strategy, Denote the model approximation loss.
[0041] Compared with the prior art, the beneficial effects of the present invention are:
[0042] 1. By making the radar confront the first jammer with multiple known jamming strategies, the present invention constructs a prior information library to establish the prior information available to the radar. Different jamming strategies will form different environmental dynamic mechanisms, and the commonly used jamming strategies can be used as the prior information of the FA radar to accelerate the learning of the radar anti-jamming strategy. Therefore, the present invention applies the expert knowledge about the jamming strategy to the radar to avoid the radar learning from scratch and improve the efficiency of the radar learning the anti-jamming strategy.
[0043] 2. The present invention decomposes the unknown jamming strategy into the weighted sum of the known jamming strategies in the prior information library. The maximum anti-jamming performance against the unknown jamming strategy is equivalent to the maximum anti-jamming performance of the weighted combination of the known jamming strategies, and it is used as the objective function of the radar decision-making. Only when the weight coefficient and the learning model are both optimal can the optimal solution of the objective function be obtained; further, by combining the objective function of the radar decision-making with the weight coefficient, a two-layer optimization cost function is proposed, that is, simultaneously optimizing the weight coefficient and the objective function of the radar decision-making. By solving this two-layer optimization cost function, an anti-jamming strategy with theoretical boundary guarantee can be obtained, so that the radar can learn an effective anti-jamming strategy.
[0044] 3. The present invention uses a model-based reinforcement learning method to obtain first interaction information by having a radar confront a first jammer with known multiple jamming strategies, interact with a second jammer with unknown jamming strategies to obtain second interaction information, and update the parameters of the learning model using the first interaction information and the second interaction information. Thus, radar strategy calculation is performed using the updated learning model. This method has higher sample efficiency, can construct a simulation environment based on some interaction samples, and then cheaply generate a large number of samples without interacting with the real environment. It can learn effective anti-jamming strategies using fewer interaction samples, improving the efficiency of the radar in learning anti-jamming strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 FIG. is a schematic flowchart of a radar anti-jamming strategy learning method based on knowledge and model-based reinforcement learning provided by an embodiment of the present invention;
[0046] Figure 2 FIG. is a specific flowchart of the radar anti-jamming strategy learning method based on knowledge and model-based reinforcement learning provided by an embodiment of the present invention;
[0047] Figure 3 FIG. is a schematic diagram of two storage modes of a jammer provided by an embodiment of the present invention;
[0048] Figures 4a - 4d FIG. is a schematic diagram of the anti-jamming performance of a model-free reinforcement learning method against four different jamming strategies provided by an embodiment of the present invention;
[0049] Figures 5a - 5b FIG. is a schematic diagram of the convergence detection probability per round under different interaction sample sizes provided by an embodiment of the present invention;
[0050] Figures 6a - 6d FIG. is a schematic diagram comparing the convergence detection probability of the method of this embodiment with the model-free reinforcement learning method (PPO) when the unknown jamming strategy is included in the radar prior information library provided by an embodiment of the present invention;
[0051] Figures 7a - 7d FIG. is a schematic diagram of the weight value of λ during the learning process when the unknown jamming strategy is included in the radar prior information provided by an embodiment of the present invention;
[0052] Figures 8a - 8d FIG. is a schematic diagram comparing the convergence detection probability of the method of this embodiment with the model-free reinforcement learning method (PPO) when the unknown jamming strategy is not included in the radar prior information library provided by an embodiment of the present invention;
[0053] Figures 9a - 9d FIG. is a schematic diagram of the weight value of λ during the learning process when the unknown jamming strategy is not included in the radar prior information. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The present invention will be further described in detail below with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0055] Embodiment 1
[0056] Please refer to Figure 1 and Figure 2 , Figure 1 which is a schematic flowchart of a radar anti-jamming strategy learning method based on knowledge and model-based reinforcement learning provided by an embodiment of the present invention, Figure 2 which is a specific flowchart of the radar anti-jamming strategy learning method based on knowledge and model-based reinforcement learning provided by an embodiment of the present invention.
[0057] The radar anti-jamming strategy learning method based on knowledge and model-based reinforcement learning includes the steps:
[0058] S1. Make the radar confront a first jammer with multiple known jamming strategies to perform anti-jamming strategy learning, and construct a prior information library by using the first interaction information and the learned anti-jamming strategies.
[0059] Specifically, in the radar anti-jamming strategy learning problem, the prior information includes the jamming strategies that the jammer may adopt, the interaction information between the radar and the jammer, and the anti-jamming strategies learned by the radar. For example, one of the most common jamming strategies of the jammer is to intercept part of the pulses, and then emit a targeted jamming signal according to the center frequency of the intercepted pulses; the radar interacts with the jammer that intercepts part of the pulses, and during the interaction, the radar uses a neural network to learn the jamming strategy of the jammer, so as to obtain the interaction information; the radar obtains the anti-jamming strategy after learning. Further, the jamming strategies of the jammer, the interaction information between the radar and the jammer, and the anti-jamming strategies learned by the radar are constructed into a prior information library.
[0060] Before the radar learns the anti-jamming strategy, some prior information about the jammer is available. Assume that there are D known jamming strategies that can be used to build the prior information library of the radar, which can be denoted as where represents the prior information constructed by the d-th known jamming strategy, S represents the state set, A represents the action set, R represents the reward function, and P d represents the transition probability. In this embodiment, the radar anti-jamming strategy learning is modeled as a Markov decision (MDP) process. Therefore, different jamming strategies will result in different transition probabilities. Further, based on the known prior information it can assist the radar in learning the anti-jamming strategy.
[0061] S2. Enable the radar to interact with the second jammer with an unknown jamming strategy according to the current strategy to collect real experience, and obtain the second interaction information between the radar and the second jammer.
[0062] Specifically, as a deep reinforcement learning agent, the radar uses a neural network to learn the jamming strategy of the jammer. The anti-jamming strategy π learned by the radar is parameterized by the neural network parameters. When facing an unknown jamming strategy, the radar interacts with the second jammer with the unknown jamming strategy according to the anti-jamming strategy π to collect a fixed number of real experiences.
[0063] In a specific embodiment, the second jammer with an unknown jamming strategy is in a time-division transceiver mode, and the duration of the basic interference division unit is T j , within the duration of one interference basic unit, the jammer can choose between two actions: intercepting or transmitting a jamming signal. There are N j interference basic units within the duration of one pulse. is the center frequency of the jamming signal of the N j th interference unit (if the action of this unit is to transmit a jamming signal). Then, the jamming action at time t can be expressed as:
[0064]
[0065] where, satisfying T r is the pulse duration, and f num is the selectable frequency point of the radar.
[0066] Furthermore, store the second interaction information between the radar and the second jammer in the memory pool. The second interaction information stored in the memory pool is:
[0067]
[0068] where, represents the memory pool for storing the second interaction information, represents the state information of the sample, represents the action taken by the intelligent radar in the current state, represents the reward received by the intelligent radar after taking the action in the current state, represents the next state reached by the intelligent radar after taking the action in the current state, N inter represents the number of samples collected, and M represents the number of pulses within one CPI.
[0069] S3. Update the parameters of the learning model using the first interaction information and the second interaction information.
[0070] Specifically, the learning model can adopt a neural network, which is used to evaluate the similarity between an unknown interference strategy and the known interference strategies in the prior information library.
[0071] The problem of the learning model for learning can be regarded as a supervised learning task. During the supervised learning process, each known prior information is used to generate a sufficient number of training samples offline to train the learning model. Specifically, each known prior information generates a sufficient number of training samples denoted as:
[0072]
[0073] where N d is the number of training samples.
[0074] Furthermore, for the problem of updating the parameters of the learning model, it can be modeled as a regression task, that is, during the process of training the learning model with the training samples generated by each known prior information the objective function of the learning model is minimized by means of stochastic gradient descent to update the parameters of the learning model. Specifically, the objective function of the learning model is:
[0075]
[0076] where, φ d represents the parameters of the learning model, represents the training samples generated by each known prior information and represents the learning model. expresses the learning model.
[0077] S5. Use the weight coefficients to decompose the unknown interference strategy into a weighted sum of the known interference strategies in the prior information library, and construct the objective function for radar decision-making.
[0078] First, use the weight coefficients to decompose the unknown interference strategy into a weighted sum of the known interference strategies in the prior information library.
[0079] Specifically, for the reinforcement learning task, learning from scratch requires a large number of interaction samples to obtain acceptable results. A simple way to reduce the sample complexity is to apply the constructed prior information to the agent. Given an unknown interference strategy this embodiment decomposes it into several parts, where each part is represented by a known interference strategy from the prior information. Mathematically, it can be expressed as:
[0080]
[0081] where, Represents an unknown interference strategy, represents the d-th known interference strategy, d = 1, 2,..., D, λ d ≥0, d = 1, 2,..., D and λ d represents the weight of each known interference strategy.
[0082] Then, based on the prior information library of the radar, the objective function of radar decision-making is constructed by combining the decomposition formula of the unknown interference strategy.
[0083] Specifically, the interference strategy determines the transition probability in the MDP, so the transition probability of can be expressed as the weighted sum of d = 1, 2,..., D. Based on the decomposition formula of the interference strategy and the prior information of the radar, the objective function of radar decision-making is:
[0084]
[0085] where, π * represents the optimal radar anti-jamming strategy, π represents the radar anti-jamming strategy, represents the state value function of the intelligent radar under the d-th known interference, and s represents the state information.
[0086] Furthermore, the goal of the radar is to maximize the above objective function.
[0087] S4. Use the updated learning model to measure the KL distance between the transition probabilities caused by the unknown interference strategy and the known interference strategies in the prior information library to evaluate the model approximation loss.
[0088] Specifically, in order to calculate the weight coefficient λ d , the radar needs to evaluate the similarity between the unknown interference strategy and the known interference strategies in the prior information library based on the real interaction samples. The similarity between the unknown interference strategy and the known interference strategy can be measured by the KL distance between the transition probabilities.
[0089] Therefore, first use the updated learning model to calculate the transition probability P(·|s, a) caused by the unknown interference strategy and the transition probability P d (·|s, a) caused by the known interference strategy. Then calculate the model approximation loss using the KL distance between the transition probability caused by the unknown interference strategy and the transition probability caused by the known interference strategy:
[0090]
[0091] where, Denote the model approximation loss, Denote the d-th known interference strategy, where d = 1, 2, ..., D, Denote the unknown interference strategy The sample distribution determined by the transition probability of and the strategy adopted by the radar, D KL Denote the KL distance between the transition probability caused by the unknown interference strategy and the transition probability caused by the known interference strategy. P(·|s,a) denotes the transition probability caused by the known interference strategy, P d (δ|s,a) denotes the transition probability caused by the known interference strategy.
[0092] S6. Evaluate the similarity between the unknown interference strategy and the known interference strategy using the model approximation loss, and calculate the weight coefficient.
[0093] Specifically, evaluate the similarity between the unknown interference strategy and the known interference strategy using the model approximation loss, and calculate the weight coefficient by solving the objective optimization problem. Therefore, based on the model approximation loss, the weight coefficient can be calculated by solving the following objective optimization problem:
[0094]
[0095] where λ * Denote the optimal weight, λ denotes the weight coefficient, λ d Denote the weight of each known interference strategy, Denote the model approximation loss.
[0096] S7. Calculate the radar anti-jamming strategy using the weight coefficient and the objective function and update it.
[0097] Specifically, after calculating the optimal weight λ * of the current true interaction information, substituting the optimal weight λ * into the objective function of the radar decision-making can obtain the optimal radar anti-jamming strategy corresponding to the current true interaction information, and update the radar strategy to this optimal radar anti-jamming strategy.
[0098] S8. Loop steps S2 - S7 until the radar performance converges or meets the preset requirements to obtain the optimal radar anti-jamming strategy.
[0099] Specifically, the convergence of the radar performance can be that the curve of the radar learning fluctuates little and the learning effect is relatively stable, etc.; the preset requirements can be that the detection probability of the radar is greater than the target probability or the duration of the radar being jammed is less than the target duration, etc. The convergence of the radar performance or meeting the preset requirements can be determined according to the actual design requirements, which will not be elaborated in this embodiment.
[0100] The objective function of the radar decision-making in this embodiment is a reinforcement learning model based on a two-layer model, and its solution is quite difficult. The main reason is that it is a nested optimization problem, and the variables to be solved in the objective function and constraints are coupled. To make the above problem feasible, this embodiment proposes an online-offline hybrid solution method based on the Dyna architecture. This method can be divided into an online part and an offline part. The online part mainly involves step S2, where the radar interacts with an unknown jammer to collect a fixed number of real experiences. The offline part involves steps S3 - S6, where the offline part uses the collected real experiences to calculate the weight vector, update the prior information, and improve the radar strategy. The online part and the offline part continuously iterate to guide the radar performance to converge or meet the predetermined requirements.
[0101] In this embodiment, the radar is made to confront a first jammer with multiple known jamming strategies, thereby constructing a prior information library and establishing the prior information available to the radar. Different jamming strategies will form different environmental dynamic mechanisms, and the commonly used jamming strategies can be used as the prior information of the FA radar to accelerate the learning of the radar anti-jamming strategy. Therefore, this embodiment applies the expert knowledge about the jamming strategy to the radar to avoid the radar learning from scratch and improves the efficiency of the radar learning the anti-jamming strategy.
[0102] This embodiment decomposes the unknown jamming strategy into a weighted sum of the known jamming strategies in the prior information library. The anti-jamming performance maximized for the unknown jamming strategy is equivalent to the anti-jamming performance maximized for the weighted combination of the known jamming strategies, and this is used as the objective function of the radar decision-making. Only when the weight coefficient and the learning model are both optimal can the optimal solution of the objective function be obtained. Further, by combining the objective function of the radar decision-making with the weight coefficient, a two-layer optimization cost function is proposed, that is, optimizing both the weight coefficient and the objective function of the radar decision-making simultaneously. By solving this two-layer optimization cost function, an anti-jamming strategy with theoretical boundary guarantee can be obtained, so that the radar can learn an effective anti-jamming strategy.
[0103] This embodiment uses the model-based reinforcement learning method. The first interaction information is obtained through the confrontation between the radar and the first jammer with multiple known jamming strategies, and the second interaction information is obtained by interacting with the second jammer with an unknown jamming strategy. The parameters of the learning model are updated using the first interaction information and the second interaction information, and then the updated learning model is used for radar strategy calculation. This method has higher sample efficiency. A simulation environment can be constructed based on some interaction samples, and then a large number of samples can be cheaply generated without interacting with the real environment. An effective anti-jamming strategy can be learned using fewer interaction samples, improving the efficiency of the radar learning the anti-jamming strategy.
[0104] The following further verifies and illustrates the effect of the present invention through simulation experiments.
[0105] (1) Simulation conditions:
[0106] One CPI of the radar contains 16 pulses, one pulse contains 4 sub - pulses, the selectable frequency points are 3, and the frequency hopping interval is 100 MHz.
[0107] The jammer has two possible modes, namely the interception mode and the emission mode. When the jammer works in the interception mode, it can acquire the carrier frequency of the radar, and the jammer can transmit interference signals according to the information stored in the memory. The present invention considers two storage modes of the jammer, as Figure 3 shown Figure 3 is a schematic diagram of two storage modes of the jammer provided by the embodiment of the present invention. In storage mode 1, the jammer can store the information intercepted by each radar sub - pulse in different memories. When the jammer decides to interfere with a given sub - pulse, its center frequency can be selected from the corresponding memory. Based on the actions of the radar and the jammer, the jammer has four different memories to store the corresponding intercepted information of each sub - pulse (the four different memories are distinguished by different colors). If the action of the jammer is to send an interference signal for a given sub - pulse, the collected information of this memory is empty, denoted as In storage mode 2, there is only one memory for storing the intercepted information. The jammer stores all the intercepted information in the same memory and sends interference signals based on this.
[0108] This embodiment uses common interference strategies as prior information of the FA radar to accelerate the learning of the radar anti - interference strategy. For the jammer, the controllable functions include the storage mode, the actions of interception and emission, and the memory length of the memory. Therefore, different combinations of them can be used to describe the strategy of the jammer. Some interference strategies of the jammer used in this embodiment are given in the following table.
[0109] Table 1 Prior information used by the radar
[0110] Interference strategy 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 Storage mode 1 1 1 1 1 1 2 2 2 2 2 2 2 2 2 Memory length 1 1 2 2 3 3 1 1 1 2 2 2 3 3 3 Interception model 1 2 1 2 1 2 3 4 5 3 4 5 3 4 5
[0111] (2) Contents of the simulation experiment
[0112] Figures 4a - 4d is a schematic diagram of the anti - interference performance of the model - free reinforcement learning method (model free RL) provided by the embodiment of the present invention for four different interference strategies. Here, take the learning performance of the radar against interference strategies 1, 5, 7, and 13 as an example. Among them, Figure 4a represents interference strategy 1, Figure 4b represents interference strategy 5, Figure 4c represents interference strategy 7, Figure 4d represents interference strategy 13. In Figures 4a - 4dAmong them, the horizontal axis represents the number of training rounds, and the vertical axis represents the radar detection probability. Compared with the strategy of randomly selecting actions, the radar based on the model-free reinforcement learning method can learn an effective anti-jamming strategy. However, at least 1×10 5 samples are required to achieve acceptable detection performance.
[0113] Figures 5a - 5b It is a schematic diagram of the convergence detection probability of each round under different interaction sample sizes provided by the embodiments of the present invention. Figure 5a represents the jamming strategy of jammer storage mode 1, storage length 1, and interception mode {(1, 1, 0, 0), (0, 0, 1, 1)}. Figure 5b represents the jamming strategy of jammer storage mode 2, storage length 2, and interception mode {(1, 1, 0, 0), (0, 0, 1, 1)}. In this embodiment, an online-offline hybrid method is adopted, that is, the radar first interacts with the jammer online to collect a fixed number of real samples, and then updates the model and improves the strategy in an offline manner. In the above process, the performance of this method is related to the online sample size of the radar and the jammer in each round. A large number of real interaction sample sizes can obtain a very accurate model and a high detection probability. However, it may take a long time to collect enough real interaction samples. In Figures 5a - 5b Among them, the horizontal axis represents the online-offline interaction rounds, and the vertical axis represents the radar detection probability. The total sum of the actual interaction sample sizes in each round can be multiplied by 2000, 10000, or 50000. As can be seen from Figures 5a - 5b , when the actual interaction sample size is 2000, good detection performance can already be obtained, and continuing to increase the actual interaction sample size cannot significantly improve the performance of the model.
[0114] Figures 6a - 6d It is a schematic diagram for comparing the convergence detection probability of the method in this embodiment with the model-free reinforcement learning method (PPO) when the unknown jamming strategy is included in the radar prior information library of the embodiments of the present invention (that is, the unknown jamming strategy is exactly matched with a certain jamming strategy in the prior information). The horizontal axis represents the actual interaction sample size before the current round, and the vertical axis represents the radar detection probability. Figure 6a represents jamming strategy 1. Figure 6b represents jamming strategy 5. Figure 6c represents jamming strategy 7. Figure 6d represents jamming strategy 13. As shown in Figure 6, in the case of the same and small number of real interaction samples, the performance of the method proposed in this embodiment is significantly better than the other two methods and is close to the best performance.
[0115] Figures 7a - 7d It is a schematic diagram of the weight of λ during the learning process when the radar prior information in the embodiments of the present invention includes an unknown jamming strategy. In Figure 7, the horizontal axis represents the prior information serial number, and the vertical axis represents the weight of λ.Figure 7a Denotes interference strategy 1, Figure 7b Denotes interference strategy 5, Figure 7c Denotes interference strategy 7, Figure 7d Denotes interference strategy 13. Obviously, λ is a one - hot encoding and "1" only appears at the prior information corresponding to the unknown interference strategy. It can be seen from the above simulation results that the existence of prior information enables the radar to improve the strategy based on accurate simulation experience, thus making the proposed method have sample efficiency.
[0116] Figures 8a - 8d This is a schematic diagram for comparing the convergence detection probability of the method in this embodiment with the model - free reinforcement learning method (PPO) when the unknown interference strategy provided by the embodiment of the present invention is not included in the radar prior information library (that is, the unknown interference strategy does not completely match all the interference strategies in the prior information). Four unknown interference strategies are given in Table 2. Among them, interference strategies 16, 17, and 18 are interception modes unknown in the prior information, while the storage length and interception method of interference strategy 19 are not shown in the prior information. In Figures 8a - 8d it, the horizontal axis represents the actual interaction sample size, and the vertical axis represents the radar detection probability, Figure 8a Denotes interference strategy 16, Figure 8b Denotes interference strategy 17, Figure 8c Denotes interference strategy 18, Figure 8d Denotes interference strategy 19. As Figures 8a - 8d shown, even if the unknown interference strategy is not included in the prior information, the performance of the method proposed in this embodiment is still much better than that of the model - free reinforcement learning method.
[0117] Table 2 Four unknown interference strategies not included in the prior information
[0118] Interference strategy 16 17 18 19 Storage mode 1 1 2 1 Memory length 1 2 2 4 Interception model {(1,0,1,0)} {(1,0,1,0),(0,1,0,1)} {(1,1,0,0),(0,0,1,1)} {(1,0,1,0)}
[0119] Figures 9a - 9d This is a schematic diagram of the weight value of λ during the learning process when the unknown interference strategy is not included in the radar prior information. λ is no longer a one - hot encoding vector because the unknown interference strategy cannot be completely expressed by a single interference strategy involved in the prior information. There are several elements greater than zero in λ, and these elements will not remain fixed during the learning process. In Figures 9a - 9d it, the horizontal axis represents the prior information serial number, and the vertical axis represents the weight value of λ, Figure 9a Denotes interference strategy 16, Figure 9b Denotes interference strategy 17, Figure 8c Denotes interference strategy 18, Figure 9dDenote the jamming strategy 19. For the jamming strategies 16 and 17, the differences among different elements in λ are very small at the beginning of the learning process. When more real interaction samples are collected, λ gradually becomes a vector containing one or two dominant elements, which means that the unknown jamming strategy can mainly be represented by one or two jamming strategies in the prior information. Compared with the jamming strategies 16 and 17, the final result of λ for the jamming strategies 18 and 19 does not have dominant elements, which means that most jamming strategies in the prior information are needed to represent the unknown jamming strategy. Therefore, different unknown jamming strategies have different characteristics, and each jamming strategy in the prior information plays a different role in expressing the unknown jamming strategy.
[0120] In summary, the radar anti-jamming strategy learning method based on knowledge and model-based reinforcement learning in this embodiment can enable the FA radar to learn effective anti-jamming strategies with fewer interaction samples, which is an efficient learning method.
[0121] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for learning a radar anti-jamming strategy based on knowledge and model-based reinforcement learning, characterized in that, Including the steps: S1. Make the radar confront a first jammer with known multiple jamming strategies for anti-jamming strategy learning, and construct a prior information library by using the first interaction information and the learned anti-jamming strategies; S2. Make the radar select actions according to the current strategy to interact with a second jammer with unknown jamming strategies to collect real experience, and obtain the second interaction information between the radar and the second jammer; S3. Update the parameters of the learning model by using the first interaction information and the second interaction information; S4. Decompose the unknown jamming strategy into a weighted sum of the known jamming strategies in the prior information library by using the weight coefficient, and construct the objective function for the radar decision-making; S5. Measure the KL distance between the transfer probability caused by the unknown jamming strategy and the transfer probability caused by the known jamming strategies in the prior information library by using the updated learning model to evaluate the model approximation loss; S6. Evaluate the similarity degree between the unknown jamming strategy and the known jamming strategies by using the model approximation loss, and calculate the weight coefficient; S7. Calculate the radar anti-jamming strategy by using the weight coefficient and the objective function and update it; S8. Loop steps S2 - S7 until the radar performance converges or meets the preset requirements to obtain the optimal radar anti-jamming strategy.
2. The method for learning a radar anti-jamming strategy based on knowledge and model-based reinforcement learning according to claim 1, wherein Step S2 includes: Make the radar select actions according to the current strategy to interact with a second jammer with unknown jamming strategies to collect real experience, and obtain the second interaction information between the radar and the second jammer; Store the second interaction information in the memory pool: Among them, represents the memory pool for storing the second interaction information, represents the status information of the sample, represents the action taken by the intelligent radar in the current state, represents the reward received by the intelligent radar after taking an action in the current state, represents the next state reached by the intelligent radar after taking an action in the current state, N inter represents the number of samples collected, and M represents the number of pulses within one CPI.
3. The method for learning a radar anti-jamming strategy based on knowledge and model-based reinforcement learning according to claim 1, wherein Step S3 includes: Use the first interaction information and the second interaction information to minimize the objective function of the learning model by means of stochastic gradient descent to update the parameters of the learning model, where the objective function of the learning model is: Among them, φ d represents the network parameters of the learning model, represents each known prior information to generate training samples, expresses the learning model.
4. The method for learning a radar anti-jamming strategy based on knowledge and model-based reinforcement learning according to claim 1, characterized in that, Step S4 includes: Decompose the unknown jamming strategy into a weighted sum of the known jamming strategies in the prior information library by using the weight coefficient: Among them, represents an unknown interference strategy, represents the d-th known interference strategy, d = 1, 2,..., D, λ d ≥ 0, d = 1, 2,..., D and λ d represents the weight of each known interference strategy; Based on the prior information library of the radar, construct the objective function for the radar decision-making in combination with the decomposition formula of the unknown jamming strategy: Among them, π * represents the optimal radar anti-jamming strategy, and π represents the radar anti-jamming strategy. represents the state value function of the intelligent radar under the d-th known interference, and s represents the state information.
5. The method for learning a radar anti-jamming strategy based on knowledge and model-based reinforcement learning according to claim 1, characterized in that Step S5 includes: Calculate the transfer probability caused by the unknown jamming strategy and the transfer probability caused by the known jamming strategies by using the updated learning model; Calculate the model approximation loss by using the KL distance between the transfer probability caused by the unknown jamming strategy and the transfer probability caused by the known jamming strategies.
6. The method for learning a radar anti-jamming strategy based on knowledge and model-based reinforcement learning according to claim 1, characterized in that The model approximation loss is: Among them, represents the model approximation loss, represents the d-th known interference strategy, where d = 1, 2,..., D, represents the unknown interference strategy is the sample distribution determined by the transition probability of KL the unknown interference strategy and the strategy adopted by the radar, and D d (·|s,a) represents the KL distance between the transition probability caused by the unknown interference strategy and the transition probability caused by the known interference strategy. P(·|s,a) represents the transition probability caused by the known interference strategy, and P 7. The method for learning a radar anti-jamming strategy based on knowledge and model-based reinforcement learning according to claim 1, wherein Step S6 includes: Evaluate the similarity degree between the unknown jamming strategy and the known jamming strategies by using the model approximation loss, and calculate the weight coefficient by solving the objective optimization problem, where the objective optimization problem is: Among them, λ * represents the optimal weight, λ represents the weight coefficient, and λ d represents the weight of each known interference strategy, represents the model approximation loss.
Citation Information
Patent Citations
Radar anti-interference rapid decision-making system and method capable of sharing migratable multi-scene features
CN111707993A
Method for generating radar intelligent cognitive anti-interference strategy
CN112904290A