Radar anti-interference strategy generation method based on model-dependent reinforcement learning
Through a model-dependent reinforcement learning method, real interaction trajectory data and virtual jammer model are used to generate virtual anti-interference strategies, optimize radar anti-interference strategies, solve the problem of low sample efficiency, and improve the radar's anti-interference ability and adaptability.
Patent Information
- Application Number
- CN202510509159.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In the prior art, deep reinforcement learning has low sample efficiency in the acquisition of radar anti-interference strategies, which is difficult to meet the real-time anti-interference needs, and the training process is time-consuming and costly.
By obtaining real interaction trajectory data for model-free reinforcement learning training, we can determine whether to start the virtual jammer model training strategy, use the virtual jammer model to generate a virtual anti-jammer strategy, and fight against the real jammer to optimize the final anti-jammer strategy.
The sample efficiency obtained by radar anti-jamming strategy is improved, the robustness and adaptability of anti-jamming strategy is improved, the dependence on real interactive data is reduced, and the radar's anti-jamming ability is enhanced.
Smart Images

Figure CN120405579A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of radar communication technology, and in particular to a radar anti-interference strategy generation method based on model-dependent reinforcement learning. Background Art
[0002] In modern electronic warfare environments, with the rapid development of advanced jamming technologies such as Digital Radio Frequency Memory (DRFM), the electromagnetic interference environments faced by radars are becoming increasingly complex and variable. These jamming technologies can generate highly realistic deceptive or targeted signals, severely impacting radar detection performance and survivability. Traditional radar anti-jamming strategies often rely on manual pre-setting, whereby specific anti-jamming measures are pre-designed based on known jamming types and parameters. However, when faced with complex and variable unknown jamming, this manual pre-setting approach proves insufficient, making it difficult to quickly and accurately identify jamming types and adjust anti-jamming strategies. Therefore, improving radar's autonomous perception and intelligent decision-making capabilities in complex electromagnetic environments has become a key issue that needs to be addressed in current radar technology.
[0003] To address these challenges, deep reinforcement learning technology has been introduced into the research of radar anti-interference strategies in recent years, achieving significant progress. As a key branch of machine learning, deep reinforcement learning builds an interaction model between an agent and its environment, enabling the agent to continuously optimize its behavior based on environmental feedback to maximize long-term benefits. In the field of radar anti-interference, deep reinforcement learning technology is used to train agents, enabling them to autonomously perceive interference strategies from radar transmission signals and echo signals containing interference components, and generate corresponding anti-interference strategies. Specifically for frequency-agile radars, deep reinforcement learning technology successfully induces jammers to interfere with the wrong frequency band by optimizing carrier frequency selection strategies, thereby improving the separability of radar and interference signals in the time-frequency domain and enhancing the radar's anti-interference capability.
[0004] While deep reinforcement learning-based radar anti-jamming strategy acquisition methods have shown great potential, existing technologies still face several pressing challenges. Specifically, low sample efficiency is one of the main bottlenecks restricting the application of deep reinforcement learning in radar anti-jamming. Deep reinforcement learning algorithms require a large amount of interaction data to train agent policies. In practical applications, radars often struggle to obtain sufficient valid echo signals from moving targets to support policy learning. This results in a time-consuming and costly training process, and makes it difficult to meet the requirements of real-time anti-jamming. Therefore, improving the sample efficiency of deep reinforcement learning in radar anti-jamming strategy acquisition has become a key issue that needs to be addressed. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a method for generating a radar anti-jamming strategy based on model-dependent reinforcement learning.
[0006] The technical problems to be solved by the present invention are realized through the following technical solutions:
[0007] In a first aspect, the present invention provides a method for generating a radar anti-jamming strategy based on model-dependent reinforcement learning, including:
[0008] Obtain the current real interaction trajectory data; the current real interaction trajectory data is the interaction data during the online confrontation between the radar and the real jammer.
[0009] Use the radar-side data in the current real interaction trajectory data to perform model-free reinforcement learning (MFRL) training to obtain the current real anti-jamming strategy.
[0010] Based on the current real anti-jamming strategy, determine whether to start the virtual jammer model training strategy.
[0011] When the virtual jammer model training strategy is started, use the jammer-side data in the current real interaction trajectory data to train a preset supervised model to obtain a virtual jammer model.
[0012] Generate a virtual anti-jamming strategy through the virtual jammer model and the radar.
[0013] Confront the virtual anti-jamming strategy with the real jammer to generate an adversarial interaction trajectory.
[0014] Determine the final anti-jamming strategy based on the adversarial interaction trajectory and the nearest current real interaction trajectory data; the final anti-jamming strategy is used to distinguish target signals and interference signals in the radar echo signal.
[0015] Optionally, the current real interaction trajectory data includes: radar state, radar action, jammer state, jammer action, and radar anti-jamming ability reward.
[0016] The radar state, radar action, and radar anti-jamming ability reward are the radar-side data in the current real interaction trajectory data.
[0017] The jammer state and jammer action are the jammer-side data in the current real interaction trajectory data.
[0018] Optionally, determining whether to start the virtual jammer model training strategy based on the current real anti-jamming strategy includes:
[0019] When the preset start condition is satisfied, start the virtual jammer model training strategy.
[0020] When the preset startup condition is not met, the virtual jammer model training strategy is not started, and the current real anti-jamming strategy is directly adopted as the final anti-jamming strategy;
[0021] The preset startup condition is judged based on the loss value of the current real anti-jamming strategy.
[0022] Optionally, the preset startup condition includes: the number of times of MFRL training meets the preset iteration times and the loss persistence of the current real anti-jamming strategy is greater than the preset first loss threshold.
[0023] Optionally, before generating the virtual anti-jamming strategy through the virtual jammer model and the radar, it further includes:
[0024] When the loss value of the virtual jammer model continuously exceeds the preset second loss threshold, the current real anti-jamming strategy is directly adopted as the final anti-jamming strategy.
[0025] Optionally, generating the virtual anti-jamming strategy through the virtual jammer model and the radar includes:
[0026] When the loss value of the virtual jammer model continuously is less than the preset second loss threshold, the virtual jammer model is used to interact with the radar to generate virtual interaction trajectory data;
[0027] The model-free reinforcement learning MFRL training is carried out using the radar-side data in the virtual interaction trajectory data to obtain the virtual anti-jamming strategy.
[0028] Optionally, determining the final anti-jamming strategy according to the adversarial interaction trajectory and the nearest current real interaction trajectory data includes:
[0029] Comparing the radar anti-jamming ability reward in the adversarial interaction trajectory with the radar anti-jamming ability reward in the nearest current real interaction trajectory data;
[0030] When the radar anti-jamming ability reward in the adversarial interaction trajectory is greater than the radar anti-jamming ability reward in the nearest current real interaction trajectory data, the virtual anti-jamming strategy is used as the final anti-jamming strategy;
[0031] When the radar anti-jamming ability reward in the adversarial interaction trajectory is less than the radar anti-jamming ability reward in the nearest current real interaction trajectory data, the current real anti-jamming strategy is used as the final anti-jamming strategy.
[0032] In a second aspect, the present invention provides a radar anti-jamming strategy generation device based on model-dependent reinforcement learning. The radar anti-jamming strategy generation device based on model-dependent reinforcement learning includes: an acquisition unit, a training unit, a judgment unit, a generation unit, an adversarial unit, and a determination unit;
[0033] The acquisition unit is used to: acquire the current real interaction trajectory data; the current real interaction trajectory data is the interaction data during the online confrontation between the radar and the real jammer;
[0034] The training unit is used to: perform model-free reinforcement learning (MFRL) training using the radar-side data in the current real interaction trajectory data to obtain the current real anti-jamming strategy;
[0035] The judgment unit is used to: judge whether to start the virtual jammer model training strategy based on the current real anti-jamming strategy;
[0036] The training unit is also used to: when the virtual jammer model training strategy is started, train a preset supervised model using the jammer-side data in the current real interaction trajectory data to obtain the virtual jammer model;
[0037] The generation unit is used to: generate a virtual anti-jamming strategy through the virtual jammer model and the radar;
[0038] The confrontation unit is used to: confront the virtual anti-jamming strategy with the real jammer to generate a confrontation interaction trajectory;
[0039] The determination unit is used to: determine the final anti-jamming strategy according to the confrontation interaction trajectory and the nearest current real interaction trajectory data; the final anti-jamming strategy is used to distinguish the target signal and the interference signal in the radar echo signal.
[0040] In a third aspect, the present invention provides a radar anti-jamming strategy generation device based on model-dependent reinforcement learning, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the radar anti-jamming strategy generation device based on model-dependent reinforcement learning runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the radar anti-jamming strategy generation method based on model-dependent reinforcement learning as described in the first aspect above.
[0041] The present invention provides a method for generating a radar anti-jamming strategy based on model-dependent reinforcement learning, including: obtaining current real interaction trajectory data; the current real interaction trajectory data being the interaction data during the online confrontation between a radar and a real jammer; using the radar-side data in the current real interaction trajectory data for model-free reinforcement learning (MFRL) training to obtain the current real anti-jamming strategy; determining whether to initiate a virtual jammer model training strategy based on the current real anti-jamming strategy; when the virtual jammer model training strategy is initiated, using the jammer-side data in the current real interaction trajectory data to train a pre-set supervised model to obtain a virtual jammer model; generating a virtual anti-jamming strategy through the virtual jammer model and the radar; conducting confrontation between the virtual anti-jamming strategy and the real jammer to generate a confrontation interaction trajectory; determining the final anti-jamming strategy based on the confrontation interaction trajectory and the nearest current real interaction trajectory data; the final anti-jamming strategy being used to distinguish target signals and interference signals in radar echo signals. In the present invention, a virtual anti-jamming strategy is generated through the virtual jammer model and the radar, and the virtual anti-jamming strategy is used to confront the real jammer to generate a confrontation interaction trajectory. This process is essentially based on real interaction data, generating a large amount of virtual interaction data through the model, expanding the scale of the training data set, and reducing the dependence of existing reinforcement learning on real interaction data. Secondly, through the confrontation between the virtual anti-jamming strategy and the real jammer, the effectiveness of the virtual anti-jamming strategy can be verified, and the strategy can be further optimized according to the confrontation result, improving the robustness and adaptability of the final anti-jamming strategy. Further, according to the judgment result of the confrontation interaction trajectory and the real interaction trajectory, the final anti-jamming strategy is dynamically adjusted, enabling the final anti-jamming strategy to adapt to the changes of the jammer in the real environment and maintaining the effectiveness of the final anti-jamming strategy. Through the above improvements, the sample efficiency in obtaining the radar anti-jamming strategy is jointly improved, and the problem of low sample efficiency in existing solutions is solved.
[0042] The following will further elaborate on the present invention in detail with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic flowchart of a method for generating a radar anti-jamming strategy based on model-dependent reinforcement learning provided by an embodiment of the present invention;
[0044] Figure 2 Exemplarily shows a schematic diagram of the interaction between a frequency-agile radar and a narrowband aiming jammer;
[0045] Figure 3 Exemplarily shows a schematic diagram of the network model structure corresponding to the anti-jamming strategy;
[0046] Figure 4 Exemplarily shows a schematic diagram of the radar state slice structure;
[0047] Figure 5 Schematic diagram showing the differences between the exemplary virtual jammer model and the real jammer;
[0048] Figure 6 Schematic diagram showing the anti-jamming strategy generated based on the virtual jammer model;
[0049] Figure 7 Graph showing the comparison results of the method of the present invention and model-free reinforcement learning with respect to the learning curve;
[0050] Figure 8 Schematic diagram of the structure of a radar anti-jamming strategy generation device based on model-dependent reinforcement learning provided by an embodiment of the present invention;
[0051] Figure 9 Schematic diagram of the structure of a radar anti-jamming strategy generation device based on model-dependent reinforcement learning provided by an embodiment of the present invention. Detailed implementation manners
[0052] The following further describes the present invention in detail with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0053] In order to improve the sample efficiency in obtaining the radar anti-jamming strategy and enhance the robustness and adaptability of the final anti-jamming strategy, an embodiment of the present invention provides a radar anti-jamming strategy generation method based on model-dependent reinforcement learning. Figure 1 Schematic flowchart of a radar anti-jamming strategy generation method based on model-dependent reinforcement learning provided by an embodiment of the present invention. As Figure 1 shown, it includes:
[0054] S101. Obtain the current real interaction trajectory data.
[0055] The current real interaction trajectory data is all the interaction data until the current time point during the online confrontation between the radar and the real jammer. Optionally, the current real interaction trajectory data includes: radar state, radar action, jammer state, jammer action, and radar anti-jamming ability reward;
[0056] The radar state, radar action, and radar anti-jamming ability reward are the radar-side data in the current real interaction trajectory data;
[0057] The jammer state and jammer action are the jammer-side data in the current real interaction trajectory data.
[0058] It should be noted that the radar in this embodiment is an intra-pulse frequency-agile radar with multiple operating frequency points, and each sub-pulse selects a frequency point as the carrier frequency; the real jammer has the digital radio frequency storage function and adopts the time-division transceiver system, and emits a narrowband targeting noise suppression interference signal. The operating duration (interception duration) of the receiving end of the real jammer and the operating duration (interference duration) of the transmitting end are both integer multiples of the radar sub-pulse width, and multiple frequency points can be interfered simultaneously. The differences between different interference strategies are mainly reflected in aspects such as the interception duration, interference duration, and selection of interference frequency points of the jammer.
[0059] In this embodiment, referring to Figure 2 , Figure 2 is the interaction schematic diagram between the frequency-agile radar and the narrowband targeting interference: For the t-th radar pulse signal, the intra-pulse frequency-agile radar selects the sub-pulse carrier frequency according to the echo signal (state s t ) and the carrier frequency selection strategy π and emits the pulse signal (action a t ); the jammer first intercepts the radar emission signal and then emits the corresponding targeting interference signal; the radar receives the new echo signal (state s t+1 ) and obtains its anti-jamming ability from it (reward r t ). The above interaction process continues until the end of a coherent processing interval (CPI), and the radar updates the carrier frequency selection strategy based on the rewards obtained from each interaction. It should be noted that the radar echo signal is not the only choice to represent the agent state, and any tensor containing "radar emission signal" and "interference signal" can be used as the state input to the radar. Therefore, the "echo signal" mentioned later also refers to such tensors.
[0060] In this embodiment, a radar emission pulse is equally divided into M sub-pulses, and the carrier frequency of each sub-pulse can be arbitrarily selected from the given N frequency points. Regarding the pulse signal emitted by the radar as the agent action, the radar action expression is:
[0061]
[0062] Among them, represents the t-th pulse signal emitted by the radar, d M represents the frequency point corresponding to the M-th sub-pulse, and N represents the total number of frequency points corresponding to the M-th sub-pulse.
[0063] Since a radar pulse contains multiple sub-pulses and all sub-pulses share an optional frequency point set Therefore, the radar action set can be expressed as:
[0064]
[0065] Among them, × represents the Cartesian product.
[0066] In this embodiment, the radar echo signal contains interference components. When the transceiver time-division jammer (real jammer) emits narrowband aiming jamming, the receiving end and the transmitting end work alternately: when the receiving end works, the jammer intercepts a section of the radar transmitted signal, obtains its frequency information through waveform analysis, and stores its carrier frequency (frequency point) in the jammer memory bank; when the transmitting end works, the jammer selects a certain number of frequency points from its own memory bank according to the preset jamming strategy, and emits narrowband aiming jamming signals for the selected frequency points. The working mode of the jammer switches repeatedly between intercepting signals and implementing jamming until the receiving end can no longer receive the radar transmitted signal, and the jammer stops working. Therefore, the actions of the jammer can be expressed as:
[0067]
[0068] Among them, represents the t-th jamming pulse signal emitted by the jammer, which is a matrix of M rows and N columns. Each row represents a jamming sub-pulse, and the m-th jamming sub-pulse is used to jam the m-th radar sub-pulse; represents whether the m-th jamming sub-pulse covers the radar working frequency band corresponding to the n-th frequency point, corresponding to not covered. By analogy with the action set of the agent, the action set of the jammer is defined as If all elements in the m-th row are 0, that is:
[0069]
[0070] it means that the jammer has intercepted the m-th sub-pulse of the radar transmission. At this time, the transmitting end of the jammer pauses working.
[0071] In this embodiment, since in the real electromagnetic game scenario, the power of the interference component received by the radar is generally much greater than the power of the real target echo component, the radar can relatively easily obtain the time-frequency diagram of the interference signal using methods such as spectrum analysis. Therefore, by combining its own transmitted signal and the received interference signal, the radar can construct an "echo signal" representing the state tensor, that is, the state s when the radar emits the t-th pulse t can be represented by the past k "echo signals". Therefore, the radar state s t is expressed as:
[0072]
[0073] Among them, represents the (t - 1)-th pulse signal emitted by the radar, Denote the (t - 1)-th interference pulse signal emitted by the jammer.
[0074] In this embodiment, the jammer takes the last T actions executed by the radar as the jammer state That is:
[0075]
[0076] In this embodiment, the radar uses the unobscured rate as an index to measure its own anti-jamming ability (i.e., the radar anti-jamming ability reward), that is:
[0077]
[0078] where UCR represents the unobscured rate of the radar pulse signal, and l m ∈{0, 1} indicates whether the m-th radar sub-pulse is jammed. When l m = 0, it means that the m-th interference sub-pulse does not allocate interference power in the working frequency band corresponding to the radar carrier frequency (i.e., ), and the m-th radar sub-pulse is not jammed at all. The radar can enhance its anti-jamming ability by increasing UCR.
[0079] In addition, the current real interaction trajectory data in S101 is also obtained and analyzed based on the radar echo data.
[0080] S102. Use the radar-side data in the current real interaction trajectory data to perform model-free reinforcement learning MFRL training to obtain the current real anti-jamming strategy.
[0081] In the embodiment of the present invention, for example, the proximal policy optimization algorithm (PPO) can be used to perform model-free reinforcement learning MFRL training to obtain the current real anti-jamming strategy.
[0082] The loss function l(φ) during the model-free reinforcement learning MFRL training guided by the proximal policy optimization algorithm can be expressed as:
[0083]
[0084] where, denotes the mathematical expectation, and the state and action combination of the radar follow the distribution of the policy ; φ represents the policy network parameters of the agent (radar) at the current moment; φ old represents the initial parameters of the policy network; denotes the advantage function; ε is a truncation parameter with a value in (0, 1); clip(x, l, r) = max(min(x, r), l), which means restricting x within the interval [l, r]. By minimizing the loss function, the radar can optimize its anti-jamming strategy and obtain the current true anti-jamming strategy. Here, x represents the input parameter, l represents the assumed minimum value of the input parameter x, and r represents the assumed maximum value of the input parameter x.
[0085] It should be emphasized that the processing procedures of S101 - S102 are continuously executed. Therefore, the current true anti-jamming strategy is also continuously updated.
[0086] S103. Determine whether to initiate the training strategy of the virtual jammer model based on the current true anti-jamming strategy.
[0087] Optionally, S103 may specifically include:
[0088] When the preset activation condition is satisfied, initiate the training strategy of the virtual jammer model;
[0089] When the preset activation condition is not satisfied, do not initiate the training strategy of the virtual jammer model and directly adopt the current true anti-jamming strategy as the final anti-jamming strategy;
[0090] The preset activation condition is based on the loss value of the current true anti-jamming strategy for judgment.
[0091] Optionally, the preset activation condition includes: the number of MFRL training times meets the preset iteration times and the loss of the current true anti-jamming strategy is continuously greater than the preset first loss threshold.
[0092] S104. When the training strategy of the virtual jammer model is initiated, use the jammer-side data in the current true interaction trajectory data to train the preset supervised model to obtain the virtual jammer model.
[0093] Optionally, before S105, it further includes:
[0094] When the loss value of the virtual jammer model is continuously greater than the preset second loss threshold, directly adopt the current true anti-jamming strategy as the final anti-jamming strategy.
[0095] It should be noted that the preset supervised model can be artificially constructed according to actual needs or an existing supervised learning model can be adopted. This embodiment does not make any limitations in this regard.
[0096] S105. Generate a virtual anti-jamming strategy through the virtual jammer model and the radar.
[0097] Optionally, S105 may specifically include:
[0098] When the loss value of the virtual jammer model is continuously less than a preset second loss threshold, the virtual jammer model is used to interact with the radar to generate virtual interaction trajectory data;
[0099] The radar-side data in the virtual interaction trajectory data is used for model-free reinforcement learning (MFRL) training to obtain a virtual anti-jamming strategy.
[0100] It should be noted that in S105, the judgment between the loss value of the virtual jammer model and the preset second loss threshold needs to be carried out within a preset number of iterations.
[0101] S106: The virtual anti-jamming strategy is used to confront the real jammer to generate confrontation interaction trajectories.
[0102] S107: Based on the confrontation interaction trajectory and the nearest current real interaction trajectory data, the final anti-jamming strategy is determined.
[0103] The final anti-jamming strategy is used to distinguish target signals and interference signals in the radar echo signal.
[0104] Optionally, S107 may specifically include:
[0105] Compare the radar anti-jamming ability reward in the confrontation interaction trajectory with the radar anti-jamming ability reward in the nearest current real interaction trajectory data;
[0106] When the radar anti-jamming ability reward in the confrontation interaction trajectory is greater than the radar anti-jamming ability reward in the nearest current real interaction trajectory data, use the virtual anti-jamming strategy as the final anti-jamming strategy;
[0107] When the radar anti-jamming ability reward in the confrontation interaction trajectory is less than the radar anti-jamming ability reward in the nearest current real interaction trajectory data, use the current real anti-jamming strategy as the final anti-jamming strategy.
[0108] An embodiment of the present invention provides a method for generating a radar anti-jamming strategy based on model-dependent reinforcement learning, including: obtaining current real interaction trajectory data; the current real interaction trajectory data is the interaction data during the online confrontation between a radar and a real jammer; using the radar-side data in the current real interaction trajectory data for model-free reinforcement learning (MFRL) training to obtain the current real anti-jamming strategy; determining whether to start a virtual jammer model training strategy based on the current real anti-jamming strategy; when the virtual jammer model training strategy is started, using the jammer-side data in the current real interaction trajectory data to train a preset supervised model to obtain a virtual jammer model; generating a virtual anti-jamming strategy through the virtual jammer model and the radar; conducting confrontation between the virtual anti-jamming strategy and the real jammer to generate an adversarial interaction trajectory; determining the final anti-jamming strategy based on the adversarial interaction trajectory and the nearest current real interaction trajectory data; the final anti-jamming strategy is used to distinguish target signals and interference signals in the radar echo signal. In the present invention, a virtual anti-jamming strategy is generated through the virtual jammer model and the radar, and an adversarial interaction trajectory is generated by using the virtual anti-jamming strategy to confront the real jammer. This process is essentially based on real interaction data, generating a large number of virtual interaction data through the model, expanding the scale of the training data set, and reducing the dependence of existing reinforcement learning on real interaction data. Secondly, through the confrontation between the virtual anti-jamming strategy and the real jammer, the effectiveness of the virtual anti-jamming strategy can be verified, and the strategy can be further optimized according to the confrontation results, improving the robustness and adaptability of the final anti-jamming strategy. Further, according to the judgment result of the adversarial interaction trajectory and the real interaction trajectory, the final anti-jamming strategy is dynamically adjusted, so that the final anti-jamming strategy can adapt to the changes of the jammer in the real environment and maintain the effectiveness of the final anti-jamming strategy. Through the above improvements, the sample efficiency in obtaining the radar anti-jamming strategy is jointly improved, and the problem of low sample efficiency in the existing solution is solved.
[0109] In order to verify the effectiveness of the method for generating a radar anti-jamming strategy based on model-dependent reinforcement learning provided by the present invention, a simulation experiment was also conducted.
[0110] This embodiment includes a set of experiments:
[0111] The experiment studies the confrontation process between an intra-pulse frequency-agile radar and a narrowband aiming self-defense jammer. The parameter settings related to the radar are shown in Table 1:
[0112] Table 1 Radar-related parameter settings during the experiment
[0113]
[0114] For the same narrowband aiming jamming, the differences between different jamming strategies are mainly reflected in aspects such as the intercept duration, jamming duration, and selection of jamming frequency points of the jammer. A jammer with DRFM function will store the intercepted radar sub-pulse frequency point information in its own memory bank. Obviously, the more information the memory bank can store, the stronger the jamming ability of the jammer. However, limited by hardware conditions, the memory bank of the jammer cannot be infinitely large. Therefore, it is experimentally assumed that the jammer will empty the memory bank once at the end of each CPI. In addition, considering that the working bandwidth of the jamming sub-pulse is generally larger than that of the radar sub-pulse, and the jammer can apply jamming to multiple intercepted frequency points simultaneously. Therefore, it is experimentally assumed that the frequency band of the jamming sub-pulse will cover the last 15 frequency points stored in the memory bank, as well as 10 upper and lower adjacent frequency points each (it can also be understood that the jammer emits 15 jamming sub-pulses with different carrier frequencies simultaneously, and the bandwidth of each jamming sub-pulse is 20 times the minimum hopping interval of the radar sub-pulse carrier frequency). For the convenience of subsequent description, the transceiver mode "receive m and transmit n" of the jamming strategy is defined in the experiment as: the intercept duration of the jammer is equal to the duration of m radar sub-pulses, and the jamming duration is equal to the duration of n sub-pulses. The switching time between the two working modes of the jammer is ignored.
[0115] Based on the above assumptions, 5 jamming strategies with different transceiver modes are used in the experiment: receive 1 and transmit 1, receive 1 and transmit 4, receive 2 and transmit 3, receive 4 and transmit 6, and receive 5 and transmit 5.
[0116] Since each radar sub-pulse can arbitrarily select a frequency point from the set of available frequency points, there are 20 10 actions in the action set of the radar, and the action space is very large. The neural network cannot enumerate all the radar actions. Therefore, the experiment regards the action space of the radar as a continuous action space, and the neural network directly outputs the probability distribution of the frequency point selection of each radar sub-pulse, that is, 10 neurons in the output layer of the network correspond to 10 radar sub-pulses. Figure 3 Exemplarily shows a schematic diagram of the network model structure corresponding to the anti-jamming strategy. Further, the anti-jamming strategy here can be a virtual anti-jamming strategy or a real anti-jamming strategy. As Figure 3 shown, there are 3 hidden layers in the network model structure corresponding to the anti-jamming strategy. The first two layers are convolutional neural networks (CNNs). The number of output channels of the first layer is 9, and the number of output channels of the second layer is 18. The third layer is a fully connected network with 64 neurons. The activation function of these 3 layers is the ReLU function. The output layer is divided into two parts, each consisting of 10 neurons, corresponding to the mean and variance of the 10 radar sub-pulse frequency points. The mean and variance can be used to construct the corresponding Gaussian distribution, and then sampling is performed to obtain the radar sub-pulse frequency points. Since the sampling result may not be an integer, the sampling result needs to be discretized to obtain the exact radar action. The discretization method adopted in the experiment is as follows:
[0117]
[0118] where round(·) represents the rounding function, and f con is the continuous value obtained by sampling, and f dis is the discretized radar sub-pulse frequency point, N is the number of available radar frequency points, and in this experiment N = 20. It is not difficult to see that f dis ∈{0, 1, …, N - 1}.
[0119] To use the above policy network structure, the experiment refers to the structure of image data and uses the form of a third-order tensor to represent the environmental state. Figure 4 Exemplarily, a schematic diagram of the radar state slice is shown. Since the experiment regards the "radar - jammer" actions of the past k = 3 pulses as the environmental state. Therefore, the state "image" contains 3 layers of matrices, and each layer of matrix is constructed based on the jammer actions of the corresponding pulse. As Figure 4 shown, '0' in the matrix indicates that the interfering sub-pulse does not cover this frequency point (the jammer may be intercepting the signal); '1' indicates that the interfering sub-pulse covers this frequency point. Add '2' at the position of the radar sub-pulse frequency point in the jammer action matrix. At this time, '2' will appear in the state tensor, and '3' may appear. '2' indicates that the corresponding radar sub-pulse is not interfered, and '3' indicates that the corresponding radar sub-pulse is interfered. In summary, the numbers that may appear in the state tensor are: '0', '1', '2', '3'.
[0120] In this experiment, S102 and S105 use the PPO algorithm to online train the anti-jamming strategy, and the hyperparameters involved are shown in Table 2:
[0121] Table 2 Hyperparameters corresponding to online training of anti-jamming strategy
[0122]
[0123] In this experiment, the radar will construct an environmental model and offline train a virtual anti-jamming strategy when the training round is a multiple of 10 and the non-coverage rate is lower than 90%. If the virtual anti-jamming strategy trained offline is significantly better than the real anti-jamming strategy trained online - specifically, the non-coverage rate of the signal generated by the former is 10% higher than that of the latter, then the radar will adopt the former as the new anti-jamming strategy. The parameters related to the virtual jammer construction process are shown in Table 3:
[0124] Table 3 Parameters related to virtual jammer construction process
[0125]
[0126]
[0127] Simulation content and result analysis:
[0128] Please refer to Figure 5 , Figure 5 which exemplarily shows a schematic diagram of the difference between the virtual jammer model and the real jammer. Figure 5 Specifically, it shows a schematic diagram of the real interaction trajectory collected when the radar trains the anti-jamming strategy of receiving 1 and transmitting 4 for the virtual jammer model or the real jammer, a schematic diagram of the simulated interaction trajectory generated by the environment model, and the difference between the two. Figure 5 Figure (a) shows a schematic diagram of the real interaction trajectory collected by the radar for the real jammer, Figure 5 Figure (b) shows a schematic diagram of the simulated interaction trajectory collected by the radar for the virtual jammer model, Figure 5 Figure (c) shows a schematic diagram of the difference between the real interaction trajectory and the simulated interaction trajectory. Figure 5 The yellow color block corresponding to the number '2' in [Figure] indicates the actual working frequency point of the radar sub-pulse that is not jammed. The green color block corresponding to the number '1' indicates the frequency point that is not the actual working frequency point of the radar sub-pulse covered by the jammer sub-pulse. The red color block corresponding to the number '3' indicates the actual working frequency point of the radar sub-pulse covered by the jammer sub-pulse. By comparison, it can be seen that the virtual jamming strategy (virtual jammer model) constructed based on supervised learning is generally similar to the real jamming strategy (real jammer), with only deviations at individual sub-pulses. The reason for the deviation is that the labeled data used to construct the virtual jammer is sampled from the initial stage of the radar's online training of the anti-jamming strategy. At this time, the radar's anti-jamming ability is weak, and a large number of working frequency points are covered by the jamming signal. Therefore, the virtual jammer trained based on these data will "overestimate" the jamming ability of the real jamming strategy.
[0129] Figure 6 Exemplarily shows a schematic diagram of the anti-jamming strategy generated based on the virtual jammer model. Figure 6 Figure (a) shows the anti-jamming strategy obtained by offline training of receiving 1 and transmitting 4 for the jamming strategy generated for the virtual jammer model. Figure 6 Figure (b) shows exemplarily Figure 6 the result of the online test of the strategy corresponding to Figure (a). It can be seen from Figure 6 that although there are certain deviations between the virtual jamming strategy and the real jamming strategy, the anti-jamming strategy trained by the radar for the former with slightly stronger jamming ability can still achieve good results in the real "radar-jammer" confrontation scenario.
[0130] Figure 7 Exemplarily shows a comparison result graph of the learning curve of the method of the present invention and model-free reinforcement learning. Figure 7Figure (a) exemplarily shows the comparison result graph of the learning curves of the model-dependent reinforcement learning method of the present invention and the existing model-free reinforcement learning when facing the interference strategy of receiving 1 and transmitting 1; Figure 7 Figure (b) exemplarily shows the comparison result graph of the learning curves of the model-dependent reinforcement learning method of the present invention and the existing model-free reinforcement learning when facing the interference strategy of receiving 1 and transmitting 4; Figure 7 Figure (c) exemplarily shows the comparison result graph of the learning curves of the model-dependent reinforcement learning method of the present invention and the existing model-free reinforcement learning when facing the interference strategy of receiving 2 and transmitting 3; Figure 7 Figure (d) exemplarily shows the comparison result graph of the learning curves of the model-dependent reinforcement learning method of the present invention and the existing model-free reinforcement learning when facing the interference strategy of receiving 4 and transmitting 6; Figure 7 Figure (e) exemplarily shows the comparison result graph of the learning curves of the model-dependent reinforcement learning method of the present invention and the existing model-free reinforcement learning when facing the interference strategy of receiving 5 and transmitting 5; Figure 7 The abscissa represents the number of rounds of online interaction and training of the radar for the real interference strategy. Since it is assumed in this experiment that the speed of the radar offline training the anti-interference strategy is much faster than the speed of online collecting real interaction trajectories, the "model-dependent reinforcement learning" curve ignores the time consumed for constructing the environment model and offline training the anti-interference strategy. Obviously, as the number of online training rounds increases, the unobscured rates of the two curves are gradually increasing and slowly approaching 1 (i.e., the theoretical upper limit); the "model-dependent reinforcement learning" curve has a jump during the training process (corresponding to the transfer of the offline-trained anti-interference strategy to the real electromagnetic game scenario), and converges earlier than the "model-free reinforcement learning" curve. Therefore, it is concluded that the method provided by the present invention can significantly reduce the number of real samples required for the radar to obtain the anti-interference strategy and improve the sample efficiency compared with the existing radar anti-interference strategy acquisition method based on the conventional model-free reinforcement learning algorithm.
[0131] In summary, the present invention provides a method for obtaining a radar anti-jamming strategy based on model-dependent reinforcement learning. First, by deeply mining the behavior logic of the jammer implied in the real interaction trajectory, the method significantly improves the utilization efficiency of the radar for the real interaction trajectory and reduces the number of online training rounds required to obtain an effective anti-jamming strategy. Second, since the interaction between the radar and the virtual jammer does not affect the real electromagnetic environment, the method enables the radar to repeatedly explore the optimal anti-jamming strategy offline, thus effectively alleviating the risk of the anti-jamming strategy falling into a local optimum during the online training process and further enhancing the anti-jamming ability of the radar. In addition, during the process of obtaining an effective anti-jamming strategy, the present invention has no special constraints on the specific algorithms used to construct the virtual jammer and train the radar offline or online, and has strong compatibility with the existing work in related fields. Finally, the present invention also considers the characteristics of the simultaneous actions of the radar and the jammer and multi-round interactions, and models the jammer's actions in the form of a time-frequency matrix, making the description of the interaction process between the radar and the jammer more in line with the actual situation.
[0132] The method provided by the embodiment of the present invention can be applied to an electronic device. Specifically, the electronic device can be: a desktop computer, a portable computer, a smart mobile terminal, a server, etc., which are not limited in the embodiment of the present invention.
[0133] Based on the same inventive concept, the embodiment of the present invention also provides a radar anti-jamming strategy generation device based on model-dependent reinforcement learning. Figure 8 As shown in the structure diagram of a radar anti-jamming strategy generation device based on model-dependent reinforcement learning provided by the embodiment of the present invention, Figure 8 it includes: an acquisition unit 601, a training unit 602, a judgment unit 603, a generation unit 604, an adversarial unit 605, and a determination unit 606;
[0134] The acquisition unit 601 is used to: acquire the current real interaction trajectory data; the current real interaction trajectory data is the interaction data during the online confrontation between the radar and the real jammer;
[0135] The training unit 602 is used to: perform model-free reinforcement learning (MFRL) training using the radar-side data in the current real interaction trajectory data to obtain the current real anti-jamming strategy;
[0136] The judgment unit 603 is used to: judge whether to start the virtual jammer model training strategy based on the current real anti-jamming strategy;
[0137] The training unit 602 is also used to: when the virtual jammer model training strategy is started, train a pre-set supervised model using the jammer-side data in the current real interaction trajectory data to obtain a virtual jammer model;
[0138] The generating unit 604 is configured to: generate a virtual anti-jamming strategy through a virtual jammer model and a radar;
[0139] The countermeasure unit 605 is configured to: counteract the virtual anti-jamming strategy with a real jammer to generate a countermeasure interaction trajectory;
[0140] The determining unit 606 is configured to: determine a final anti-jamming strategy according to the countermeasure interaction trajectory and the nearest current real interaction trajectory data; the final anti-jamming strategy is used to distinguish target signals and interference signals in radar echo signals.
[0141] Figure 9 FIG. is a schematic structural diagram of a radar anti-jamming strategy generation device based on model-dependent reinforcement learning provided by an embodiment of the present invention, including: a processor 710, a storage medium 720, and a bus 730. The storage medium 720 stores machine-readable instructions executable by the processor 710. When the radar anti-jamming strategy generation device based on model-dependent reinforcement learning runs, the processor 710 communicates with the storage medium 720 through the bus 730, and the processor 710 executes the machine-readable instructions to execute the steps of the above method embodiment. The specific implementation manners and technical effects are similar and will not be described in detail here.
[0142] The storage medium may include a random access memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-Volatile Memory, NVM), such as at least one disk memory. Optionally, the storage medium may also be at least one storage device located far from the aforementioned processor.
[0143] The above-mentioned processor may be a general-purpose processor, including a central processing unit (Central Processing Unit, CPU), a network processor (Network Processor, NP), etc.; it may also be a digital signal processor (Digital Signal Processing, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field-Programmable Gate Array, FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0144] It should be noted that the terms "first", "second", etc. are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention.
[0145] In the description of this specification, the description referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0146] Although the present invention has been described in connection with various embodiments herein, however, in the process of implementing the claimed present invention, those skilled in the art can understand and achieve other variations of the above-described disclosed embodiments by viewing the drawings and the disclosure. In the description of the present invention, the term "including" does not exclude other components or steps, the term "a" or "one" does not exclude a plurality of cases, and the meaning of "a plurality" is two or more unless otherwise specifically defined. In addition, certain measures are described in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0147] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for generating a radar anti-jamming strategy based on model-dependent reinforcement learning, characterized in that Including: Obtain the current real interaction trajectory data; The current real interaction trajectory data is the interaction data during the online confrontation between the radar and the real jammer; Use the radar-side data in the current real interaction trajectory data for model-free reinforcement learning (MFRL) training to obtain the current real anti-jamming strategy; Based on the current real anti-jamming strategy, determine whether to start the virtual jammer model training strategy; When the virtual jammer model training strategy is started, use the jammer-side data in the current real interaction trajectory data to train a pre-set supervised model to obtain a virtual jammer model; Generate a virtual anti-jamming strategy through the virtual jammer model and the radar; Confront the virtual anti-jamming strategy with the real jammer to generate an adversarial interaction trajectory; Based on the adversarial interaction trajectory and the nearest current real interaction trajectory data, determine the final anti-jamming strategy; The final anti-jamming strategy is used to distinguish target signals and interference signals in the radar echo signal.
2. The method for generating a radar anti-jamming strategy based on model-dependent reinforcement learning according to claim 1, characterized in that, The current real interaction trajectory data includes: radar state, radar action, jammer state, jammer action, and radar anti-jamming ability reward; The radar state, the radar action, and the radar anti-jamming ability reward are the radar-side data in the current real interaction trajectory data; The jammer state and the jammer action are the jammer-side data in the current real interaction trajectory data.
3. The method for generating a radar anti-jamming strategy based on model-dependent reinforcement learning according to claim 1, wherein The determination of whether to start the virtual jammer model training strategy based on the current real anti-jamming strategy includes: When a pre-set start condition is satisfied, start the virtual jammer model training strategy; When the pre-set start condition is not satisfied, do not start the virtual jammer model training strategy and directly use the current real anti-jamming strategy as the final anti-jamming strategy; The pre-set start condition is based on the loss value of the current real anti-jamming strategy as the judgment basis.
4. The method for generating a radar anti-jamming strategy based on model-dependent reinforcement learning according to claim 3, wherein The pre-set start condition includes: the number of times of the MFRL training satisfies a pre-set iteration number and the loss persistence of the current real anti-jamming strategy is greater than a pre-set first loss threshold.
5. The method for generating a radar anti-jamming strategy based on model-dependent reinforcement learning according to claim 1, wherein Before generating the virtual anti-jamming strategy through the virtual jammer model and the radar, further including: When the loss value of the virtual jammer model continuously exceeds a pre-set second loss threshold, directly use the current real anti-jamming strategy as the final anti-jamming strategy.
6. The method for generating a radar anti-jamming strategy based on model-dependent reinforcement learning according to claim 5, wherein The generation of the virtual anti-jamming strategy through the virtual jammer model and the radar includes: When the loss value of the virtual jammer model continuously is less than a pre-set second loss threshold, use the virtual jammer model to interact with the radar to generate virtual interaction trajectory data; Use the radar-side data in the virtual interaction trajectory data for model-free reinforcement learning (MFRL) training to obtain the virtual anti-jamming strategy.
7. The method for generating a radar anti-jamming strategy based on model-dependent reinforcement learning according to claim 1, wherein The determination of the final anti-jamming strategy based on the adversarial interaction trajectory and the nearest current real interaction trajectory data includes: Compare the radar anti-jamming ability reward in the adversarial interaction trajectory with the radar anti-jamming ability reward in the nearest current real interaction trajectory data; When the radar anti-jamming ability reward in the adversarial interaction trajectory is greater than the radar anti-jamming ability reward in the nearest current real interaction trajectory data, use the virtual anti-jamming strategy as the final anti-jamming strategy; When the radar anti-jamming ability reward in the adversarial interaction trajectory is less than the radar anti-jamming ability reward in the nearest current real interaction trajectory data, use the current real anti-jamming strategy as the final anti-jamming strategy.
8. A radar anti-jamming strategy generation device based on model-dependent reinforcement learning, characterized in that, The radar anti-jamming strategy generation device based on model-dependent reinforcement learning includes: an acquisition unit, a training unit, a judgment unit, a generation unit, an adversarial unit, and a determination unit; The acquisition unit is used to: acquire current real interaction trajectory data; the current real interaction trajectory data is the interaction data during the online confrontation between the radar and the real jammer; The training unit is used to: perform model-free reinforcement learning (MFRL) training using the radar-side data in the current real interaction trajectory data to obtain the current real anti-jamming strategy; The judgment unit is used to: judge whether to start the virtual jammer model training strategy based on the current real anti-jamming strategy; The training unit is further used to: when the virtual jammer model training strategy is started, train a pre-set supervised model using the jammer-side data in the current real interaction trajectory data to obtain a virtual jammer model; The generation unit is used to: generate a virtual anti-jamming strategy through the virtual jammer model and the radar; The adversarial unit is used to: confront the virtual anti-jamming strategy with the real jammer to generate an adversarial interaction trajectory; The determination unit is used to: determine the final anti-jamming strategy according to the adversarial interaction trajectory and the nearest current real interaction trajectory data; the final anti-jamming strategy is used to distinguish the target signal and the interference signal in the radar echo signal.
9. A radar anti-jamming strategy generation device based on model-dependent reinforcement learning, characterized in that Comprising: A processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the radar anti-jamming strategy generation device based on model-dependent reinforcement learning runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the radar anti-jamming strategy generation method according to any one of claims 1-7.
Citation Information
Patent Citations
Method for generating radar intelligent cognitive anti-interference strategy
CN112904290A
Interference strategy sensing method based on generative adversarial imitation learning
CN116643242A
Intelligent confrontation method based on neural virtual self-gaming
CN116866895A
Methods and apparatuses for training a model based reinforcement learning model
US20240378450A1
Cited By
Complex mine radar adaptive anti-interference detection method based on reinforcement learning
CN121477158A