A Frequency Agile Radar Anti-Multi-Interference Method Based on Bandit Algorithm
Through frequency agile radar combined with Bandit Algorithm, the frequency points of the radar transmit signal are adaptively adjusted, solving the problem of weak anti-interference ability of radar in multi-jammer environments, and achieving better anti-main lobe interference effect.
Patent Information
- Application Number
- CN202310163173.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-17
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-02-17
AI Technical Summary
When facing multiple flexible and changeable interference clusters, existing radars have weak anti-interference capabilities and are unable to independently deal with dynamically changing interference environments.
Bandit Algorithm, which combines frequency agile radar with reinforcement learning, uses system parameters to build a radar anti-interference cluster model, calculates return value and regret value, introduces a multi-arm machine method to optimize the objective function, adaptively adjusts the frequency point of the radar transmit signal, and improves anti-interference ability.
Effectively avoid interference coverage frequency bands, independently select the carrier frequency of the transmit pulse, and have adaptive resistance to main lobe interference, which improves the radar's anti-interference ability in a multi-jammer environment.
Smart Images

Figure CN116087889B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer information technology for radar anti - jamming, and particularly relates to a method for a frequency - agile radar to resist multiple jammings based on the Bandit Algorithm. Background Art
[0002] With the development of electronic attack - and - defense technology, electronic warfare has become a new battlefield in modern warfare. Especially in the field of radar technology, as an important part of electronic warfare, electronic jamming technology poses a huge threat to the jammed party and greatly weakens the communication and detection capabilities of the jammed party. Relying on the development of digital frequency storage technology, modern jammers can intercept and sort radar signals, and flexibly generate various different jamming signals according to the extracted typical parameters such as amplitude, frequency, pulse width, and repetition period. Therefore, how to improve the survival ability of radar in modern warfare is an important issue in the field of electronic attack - and - defense.
[0003] In recent years, there have been great improvements in the anti - jamming ability of radars. Especially for the suppression of sidelobe jamming, there are already many effective anti - jamming methods, such as ultra - low sidelobe design, generalized sidelobe cancellation, and pattern synthesis technology, etc. However, in terms of anti - main - lobe jamming, there is still a lack of effective jamming suppression methods.
[0004] The ability of radars in the prior art to resist multiple jamming environments is weak. In the actual usage environment, there are generally multiple flexible and changeable jammer clusters, while the anti - jamming means in the prior art can only pre - set anti - jamming strategies and cannot independently cope with the changeable jammer clusters. Summary of the Invention
[0005] In order to solve the above - mentioned problems existing in the prior art, the purpose of the present invention is to provide a method for a frequency - agile radar to resist multiple jammings based on the Bandit Algorithm. By combining a frequency - agile radar with the Bandit Algorithm that has been discussed for a long time in the field of reinforcement learning, a new algorithm for the radar to cope with multiple jamming environments is realized, and the anti - jamming ability of the radar is improved.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A method for a frequency - agile radar to resist multiple jammings based on the Bandit Algorithm, comprising the following steps:
[0008] S01. Collect the initial system parameters of the radar
[0009] Collect the parameters of the initial system in the interaction process between the radar and the jammer cluster, including the number of sub - pulses of the radar, the radar bandwidth, the upper limits of the transmission powers of the radar and the jammer, and the frequencies of the radar and the jammer;
[0010] S02. Construct a radar system model
[0011] Based on the bandwidth, frequency, and power collected in step S01, construct a model of a frequency-agile radar anti-jamming cluster based on intra-pulse frequency hopping; and obtain the received signal after the online interaction between the frequency-agile radar anti-jamming cluster model and the interference cluster.
[0012] S03. Calculate the return value and regret value of the frequency-agile radar anti-jamming cluster model
[0013] Based on the received signal of the frequency-agile radar anti-jamming cluster model in step S02, calculate the signal-to-interference-plus-noise ratio as the return value of this interaction; and obtain the total regret value through the return values of each pulse round.
[0014] S04. Establish an objective function
[0015] Combining the return value and regret value obtained in step S03, and the radar anti-jamming cluster model in step S02, determine the objective function of the radar anti-jamming cluster model.
[0016] S05. Introduce multiple multi-armed bandit methods to solve the objective function
[0017] Introduce multiple multi-armed bandit methods for the objective function to optimize the objective function of the frequency-agile radar anti-jamming cluster model.
[0018] S06. Obtain the optimal radar anti-jamming cluster model
[0019] Based on the multi-armed bandit method in step S05, find the result with the smallest regret value as the optimal radar anti-jamming cluster model, and calculate the bandwidth and frequency of the optimal radar anti-jamming cluster model.
[0020] S07. Simulate and calculate the optimization effect
[0021] Simulate and calculate the anti-jamming capabilities of different multi-armed bandit methods, and compare the anti-jamming capabilities of different methods through the rate of decrease of the regret value, the rate of increase of the return value, and the final results.
[0022] Furthermore, the following content is included in step S02:
[0023] Create a frequency-agile radar anti-jamming cluster model according to the single-pulse carrier during the interaction between the radar and the interference cluster in each pulse period.
[0024] Furthermore, the following content is included in step S02:
[0025] Create a frequency-agile radar anti-jamming cluster model according to the corresponding frequencies of multiple sub-pulses in each single pulse.
[0026] Further, the step S02 includes the following: The agile radar anti-jamming cluster model includes the transmitting and receiving models of the radar and the transmitting and receiving models of the jammer.
[0027] Further, the step S02 includes the following:
[0028] The transmitting signal model of the agile radar is created according to the following function formula:
[0029]
[0030] where is the number of sub-pulses within a single pulse of the radar, is the complex envelope of the radar signal, is the sub-pulse duration of the radar, is the rectangular function, is the exponential function, is the time;
[0031] is the carrier frequency corresponding to the radar sub-pulse, is selected from the frequency set ;
[0032] Each radar sub-pulse arbitrarily selects the carrier frequency from the set to obtain the action space of the radar;
[0033] When there are 3 frequency points for each sub-pulse, .
[0034] Further, the step S02 includes the following:
[0035] The transmitting and receiving model of the jammer is created according to the following function formula:
[0036]
[0037] where is the complex envelope of the jammer signal, is the interference pulse duration, is the rectangular function, is the exponential function, is the time. is the carrier frequency corresponding to the interference pulse, which can also be selected from the frequency set ;
[0038] It is assumed that the jammer does not perform intra-pulse frequency hopping;
[0039] According to the three actions of the jammer within a single pulse: spot jamming, blanket jamming, and repeater jamming; the action space of the jammer is obtained , where are the actions corresponding to blanket jamming and repeater jamming, and the actions of spot jamming are all within .
[0040] Furthermore, the step S03 includes the following contents:
[0041] Reward value: The reward value of each monopulse radar is the sum of the reward values of multiple sub-pulses in each monopulse; define the reward value of each monopulse radar as ;
[0042] Regret value: Calculate the regret value according to the following function formula:
[0043]
[0044] where represents the reward for choosing the action as , represents the reward for choosing the action as ;
[0045] represents the gap between the cumulative reward value of the optimal action after T rounds and the cumulative reward value of the actual action in T rounds.
[0046] Furthermore, the reward value is calculated according to the common reward function: detection probability and / or jamming-to-noise ratio.
[0047] Furthermore, in the step S05, the following contents are included:
[0048] Introduce three multi-armed bandit methods to optimize the anti-jamming cluster model of agile radars;
[0049] S021, Exploration-Exploitation Algorithm Exp3 based on exponential weighting;
[0050] S0211, The radar selects the action of the radar according to the current action probability distribution ; ;
[0051] S0212, The radar receives the spectrogram of the current single pulse and obtains the reward value of the current single pulse therefrom ;
[0052] S0213, Calculate the cumulative reward value of each action of the radar: ;
[0053] S0214, The radar passes the cumulative reward value Update the action probability distribution , and the specific update formula is
[0054]
[0055] where represents that the radar has a total of actions;
[0056] S022, the Upper Confidence Bound algorithm UCB;
[0057] S0221, define the sampling mean and the upper confidence bound for each action of the radar;
[0058] S0222, the radar calculates the corresponding to each action , and selects the maximum value as the action of the current single pulse ;
[0059] S0223, the radar receives the spectrogram of the current round and obtains the return value of each single pulse from it ;
[0060] S0224, the radar counts the number of times the current action is selected , and updates the and the return of the current action through and , and the specific update formula is:
[0061]
[0062] where represents the total number of single pulses;
[0063] S023, the Thompson Sampling algorithm TS
[0064] S0231, define the distribution mean of each action of the radar as , and the variance as
[0065] S0232, each single pulse samples from its respective distribution for each action , and the radar selects the maximum value among them as the action of the current single pulse
[0066] S0233, the radar receives the spectrogram of the current single pulse and obtains the return value of the current single pulse from it
[0067] S0234, the radar updates the current action by feedback Update the mean and variance of the current action The specific update formulas are
[0068]
[0069] where is the variance of the noise
[0070] Finally, the step S07 includes the following content
[0071] Perform simulation calculations on the optimization effects according to the three interference modes of the jammer respectively
[0072] S031, Adaptive interference: The jammer makes an adaptive interference decision based on the frequency points of the previous 10 steps of the radar
[0073] S032, Timely forwarding: The jammer detects the frequency point of the first sub-pulse of the current monopulse of the radar and chooses to continuously interfere with this frequency point
[0074] S033, Fixed interference: The jammer takes actions according to the pre-designed action probability distribution
[0075] The beneficial effects of the present invention are
[0076] A frequency agile radar anti-multi-interference method based on the Bandit Algorithm. For the dynamic and non-stationary electromagnetic spectrum, a frequency agile radar is combined with the Bandit Algorithm that has been discussed for a long time in the field of reinforcement learning to effectively avoid the interference coverage frequency band and independently select the transmit pulse carrier frequency, so as to adaptively adjust the transmit signal frequency point of the agile radar, with the ability to adaptively resist main lobe interference and better cope with anti-main lobe interference; it improves the radar's ability to resist multi-interference in a multi-interference environment and makes up for the problem of weak anti-interference ability of the radar when facing a flexible and changeable interference cluster Brief Description of the Drawings
[0077] Figure 1 is a schematic diagram of the anti-interference cluster model of the agile radar in the frequency agile radar anti-multi-interference method based on the Bandit Algorithm of the present invention
[0078] Figure 2 is a schematic diagram of the average regret values before and after optimization of three different multi-armed bandits in the frequency agile radar anti-multi-interference method based on the Bandit Algorithm of the present invention when facing a timely received and transmitted interference cluster; Parctical is the average regret value before optimization
[0079] Figure 3It is a schematic diagram of the return values before and after optimization of three different multi-armed bandits in the frequency-agile radar anti-multi-interference method based on the Bandit Algorithm of the present invention when facing an interference cluster with timely transmission and reception; Parctical is the return value before optimization;
[0080] Figure 4 It is a schematic diagram of the average regret values before and after optimization of three different multi-armed bandits in the frequency-agile radar anti-multi-interference method based on the Bandit Algorithm of the present invention when facing an interference cluster with a fixed strategy; Parctical is the average regret value before optimization;
[0081] Figure 5 It is a schematic diagram of the return values before and after optimization of three different multi-armed bandits in the frequency-agile radar anti-multi-interference method based on the Bandit Algorithm of the present invention when facing an interference cluster with a fixed strategy; Parctical is the return value before optimization;
[0082] Figure 6 It is a schematic diagram of the average regret values before and after optimization of three different multi-armed bandits in the frequency-agile radar anti-multi-interference method based on the Bandit Algorithm of the present invention when facing an adaptive interference cluster; Parctical is the average regret value before optimization;
[0083] Figure 7 It is a schematic diagram of the return values before and after optimization of three different multi-armed bandits in the frequency-agile radar anti-multi-interference method based on the Bandit Algorithm of the present invention when facing an adaptive interference cluster; Parctical is the return value before optimization. Detailed implementation manners
[0084] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0085] As Figures 1 to 7 shown, to solve the problems commonly existing in the prior art, the present invention provides a frequency-agile radar anti-multi-interference method based on the Bandit Algorithm, and the overall planning scheme is as follows:
[0086] Make full use of the exploration mechanism of the Bandit Algorithm to solve the problem of agile radar against multiple interferences. For the dynamic and non-stationary electromagnetic spectrum, design an efficient exploration mechanism to adaptively adjust the transmitted signal frequency points of the agile radar, improve the anti-interference effect of the agile radar, and make up for the weak anti-interference ability of the radar in the existing technology when facing flexible and changeable interference clusters.
[0087] Greatly improve the anti-interference ability of the frequency agile radar in a multi-interference environment. In the existing technology, the radar generally resists the multi-interference environment through pre-set algorithm strategies and cannot adjust itself in real time with the change of the environment. The frequency agile radar anti-multi-interference method based on the Bandit Algorithm planned by the present invention adaptively selects and adjusts the transmitted frequency of the frequency agile radar itself in an adaptive mode with the change of the jammer and the environment, so as to obtain better anti-interference ability.
[0088] In order to better cope with main lobe interference, frequency agile technology is introduced into the field of electronic countermeasures. Compared with the traditional fixed carrier frequency radar, the intra-pulse frequency agile radar has the characteristics of effectively avoiding the interference coverage frequency band and independently selecting the transmitted pulse carrier frequency, so it has the ability to adaptively resist main lobe interference. On this basis, the algorithm planned by the present invention relies on the frequency agile radar, combines the Bandit Algorithm that has been discussed for a long time in the field of reinforcement learning, and proposes a new algorithm for the radar to cope with the multi-interference environment, greatly improving the anti-interference ability of the radar.
[0089] The specific implementation manner of the present invention is described from three aspects: scenario modeling, algorithm design and experimental effect.
[0090] I. Anti-interference cluster model of agile radar
[0091] The interaction process between the radar and the interference cluster takes a single pulse as the carrier. In each pulse time, the agile radar selects the corresponding transmitted frequency for each sub-pulse. Correspondingly, the purpose of the interference cluster is to interfere with the corresponding frequencies of each sub-pulse. A simple anti-interference model based on frequency hopping is as Figure 2 shown.
[0092] The model in the figure can be divided into two parts: the transmitting and receiving model of the radar and the corresponding transmitting and receiving model of the interference. For the agile radar, the transmitted signal can be expressed as
[0093]
[0094] where is the number of sub-pulses in a single pulse of the radar, is the complex envelope of the signal, is the sub-pulse duration, is the rectangular function, is an exponential function, where \(t\) is time. \(f_{i}\) is the carrier frequency corresponding to the \(i\)-th sub-pulse, specifically, it can be selected from the frequency set Based on this signal model, it can be found that the action space of the agile radar is related to the carrier frequency \(f_{i}\). Each sub-pulse can arbitrarily select the carrier frequency from the set Therefore, the action space of the radar Taking the case of 3 sub-pulses with 3 frequency points to choose from for each sub-pulse as an example, the total number of actions of the radar is 27.
[0095] For the jammer, similarly, the transmitted signal of the jammer is
[0096]
[0097] where \(s_{j}(t)\) is the complex envelope of the signal, \(\tau_{j}\) is the duration of the interference pulse, \(rect(t)\) is the rectangular function, \(e^{j2\pi f_{j}t}\) is the exponential function, and \(t\) is time. \(f_{j}\) is the carrier frequency corresponding to the interference pulse, which can also be selected from the frequency set Considering the interference criterion of the jammer, it is assumed here that the jammer does not perform intra-pulse frequency hopping. At the same time, within a single pulse, the jammer has three action options: spot jamming, blanket jamming, and repeater jamming. Therefore, the action space of the jammer where \(a_{2}\) and \(a_{3}\) are the actions corresponding to blanket and repeater jamming.
[0098] b) Definition of the reward. As shown above Figure 1 After the radar and the interference cluster interact in each round, the radar receives the spectrogram of the echo in the current round. The definition of the reward is based on this spectrogram. Taking 3 sub-pulses as an example, it can be observed from the final spectrogram whether the echo signal of each sub-pulse is blocked by the interference signal and the degree of blocking. Based on this, the magnitude of the reward for each sub-pulse can be obtained. And the reward of the radar in each round is the sum of the magnitudes of the rewards of the three sub-pulses.
[0099] c) Definition of the regret value. Assume that the reward value of the radar in each round is Then the regret value can be defined as
[0100]
[0101] where \(r(s,a)\) represents the reward for choosing the action as \(a\), Indicates that the selected action is of the return. It can be understood as the gap between the actual action per round and the single optimal action selected from an overall perspective. The lower the regret value, the closer the average strategy of the radar is to the optimal strategy in the extreme sense. Commonly used return functions include: detection probability, signal-to-noise ratio.
[0102] II. Types of Bandit Algorithm.
[0103] After the environment of the entire radar anti-jamming cluster is modeled, three different multi-armed bandit methods are introduced, and there are obvious improvements in the experimental results compared with the basic method.
[0104] 1) Method 1: Exploration and Exploitation Algorithm Based on Exponential Weighting (Exp3)
[0105] i) The radar selects the action of the radar according to the current action probability distribution Select the action of the radar
[0106] ii) The radar receives the spectrogram of the current round and obtains the return value of this round from it
[0107] iii) Calculate the cumulative return value of each action of the radar of:
[0108] iv) The radar updates the action probability distribution through the cumulative return value , and the specific update formula is , specifically
[0109]
[0110] where Indicates that the radar has a total of actions.
[0111] Method 2: Upper Confidence Bound Algorithm (UCB)
[0112] i) Each action of the radar has its own sampling mean and upper confidence bound
[0113] ii) The radar calculates the corresponding for each action , and selects the maximum value as the action of this round
[0114] iii) The radar receives the spectrogram of the current round and obtains the return value of this round from it
[0115] iv) The radar counts the number of selections of the current action and updates the and of the current action through and rewards. The specific update formula is: and where
[0116]
[0117] where represents the total number of rounds.
[0118] 3) Method 3: Thompson Sampling Algorithm (TS)
[0119] i) Each action of the radar has its own distribution, with a mean of and a variance of
[0120] ii) In each round, each action samples from its own distribution, and the radar selects the maximum value as the action for this round
[0121] iii) The radar receives the spectrogram of the current round and obtains the reward value for this round
[0122] iv) The radar updates the mean and variance of the current action through the reward. The specific update formula is where
[0123]
[0124] where is the variance of the noise.
[0125] Experimental results:
[0126] The purpose of the experiment is to test the radar's ability to resist interference clusters within a single pulse round. The radar is set to 3 sub-pulses, and each sub-pulse has 3 different frequency point selections. The interference cluster is set to two jammers, and the interference mode and the actions of the jammers follow the introduction in the above text. The experimental results are divided into two parts. One part is the decrease in the average regret value, and the other part is the increase in the reward value. In both parts, it can be clearly seen that after introducing the multi-armed bandit algorithm, the radar's anti-interference ability has been significantly improved. Experimental results are available for all three algorithms, represented as Exp3, UCB, and TS in the figure. The ordinary non-adaptive method is represented by Practical and is used for comparison in the figure. At the same time, three interference strategies of the jammers are defined:
[0127] i) Adaptive interference: The jammer makes an adaptive interference decision based on the frequency points of the radar in the previous 10 steps.
[0128] ii) Timely forwarding: The jammer detects the frequency point of the first sub-pulse of the radar in the current round and chooses to continuously interfere with this frequency point.
[0129] iii) Fixed interference: The jammer takes actions according to the pre-designed action probability distribution.
[0130] In the process of modeling the multi-jamming environment, the interference strategies of all jammers come from the above three strategies. Based on different interference strategies, the experimental effect diagrams are as follows.
[0131] The interference cluster adopts the timely forwarding strategy
[0132] The interference cluster adopts the fixed interference strategy
[0133] The interference cluster adopts the adaptive interference strategy
[0134] It can be seen from the above three groups of experiments that the effects of the three multi-armed bandit algorithms far exceed those of ordinary non-adaptive radar strategies. The reward value rises faster, and the average regret value also decreases faster. Among these three methods, Thompson Sampling (TS) stands out particularly.
[0135] Through the above frequency-agile radar anti-multi-jamming method based on the Bandit Algorithm, the ability of the radar to resist the interference cluster is improved. Even in the face of multiple flexible jammers, the radar can adaptively adjust its own action strategy.
[0136] Finally, a technical solution for a frequency-agile radar anti-multi-jamming method based on the Bandit Algorithm is formed: specifically, it is operated according to the following steps:
[0137] S01. Collect the initial system parameters of the radar
[0138] Collect the initial system parameters during the interaction process between the radar and the interference cluster, including the number of sub-pulses K of the radar, the radar bandwidth, the transmission powers of the radar and the jammer, where the transmission power of the radar is ; and the transmission power of the jammer , and the frequencies of the radar and the jammer;
[0139] So as to reflect the interaction process between the radar and the interference cluster by constructing a frequency-agile radar anti-jamming cluster model based on intra-pulse frequency hopping.
[0140] S02. Construct a radar system model
[0141] Construct a frequency-agile radar anti-jamming cluster model based on intra-pulse frequency hopping according to the bandwidth, frequency, and power collected in step S01; and obtain the received signal after the online interaction between the frequency-agile radar anti-jamming cluster model and the interference cluster.
[0142] Create a frequency-agile radar anti-jamming cluster model based on the monopulse carrier during the interaction between the radar and the interference cluster in each pulse period.
[0143] Create a frequency-agile radar anti-jamming cluster model according to the corresponding frequencies of multiple sub-pulses in each monopulse.
[0144] The frequency-agile radar anti-jamming cluster model includes the transmitting and receiving models of the radar and the transmitting and receiving models of the jammer.
[0145] Create the transmitting signal model of the frequency-agile radar according to the following function formula:
[0146]
[0147] Where is the number of sub-pulses within a single pulse of the radar, is the complex envelope of the radar signal, is the sub-pulse duration of the radar, is the rectangular function, is the exponential function, is the time;
[0148] is the carrier frequency corresponding to the radar sub-pulse, is selected from the frequency set ;
[0149] Each radar sub-pulse arbitrarily selects the carrier frequency from the set to obtain the action space of the radar;
[0150] and when each sub-pulse has 3 frequency points, .
[0151] And create the transmitting and receiving models of the jammer according to the following function formula:
[0152]
[0153] Where is the complex envelope of the jammer signal, is the interference pulse duration, is the rectangular function, is the exponential function, is the time. is the carrier frequency corresponding to the interference pulse, which can also be from the frequency set selected from;
[0154] Assume that the jammer does not perform intra-pulse frequency hopping;
[0155] According to the three actions of the jammer within a single pulse: spot jamming, blanket jamming, and repeater jamming; obtain the action space of the jammer , where are the actions corresponding to blanket jamming and repeater jamming, and the actions of spot jamming are all within .
[0156] S03. Calculate the reward value and regret value of the agile radar anti-jamming cluster model
[0157] Based on the received signal of the agile radar anti-jamming cluster model in step S02, calculate the signal-to-interference-plus-noise ratio as the reward value of this interaction; and further obtain the total regret value through the reward value of each pulse round;
[0158] Reward value: The reward value of each single-pulse radar is the sum of the reward values of multiple sub-pulses in each single pulse; define the reward value of each single-pulse radar as ; The reward value is calculated according to the common reward function: detection probability and / or signal-to-interference ratio.
[0159] Regret value: Calculate the regret value according to the following function formula:
[0160]
[0161] where represents the reward for choosing the action as , represents the reward for choosing the action as ;
[0162] represents the gap between the cumulative reward value of the optimal action after T rounds and the cumulative reward value of the actual action in T rounds.
[0163] S04. Establish the objective function
[0164] Combining the reward value and regret value obtained in step S03, and the radar anti-jamming cluster model in step S02, determine the objective function of the radar anti-jamming cluster model;
[0165] S05. Introduce multiple multi-armed bandit methods to solve the objective function
[0166] Introduce multiple multi-armed bandit methods for the objective function to optimize the objective function of the agile radar anti-jamming cluster model;
[0167] Specifically, introduce three multi-armed bandit methods to optimize the agile radar anti-jamming cluster model;
[0168] S021, Exploration and Exploitation Algorithm Exp3 Based on Exponential Weighting;
[0169] S0211, The radar selects the action of the radar according to the current action probability distribution ; ;
[0170] S0212, The radar receives the spectrogram of the current monopulse and obtains the return value of the current monopulse from this ;
[0171] S0213, Calculate the cumulative return value of each action of the radar : ;
[0172] S0214, The radar updates the action probability distribution through the cumulative return value , and the specific update formula is
[0173]
[0174] where indicates that the radar has a total of actions;
[0175] S022, Upper Confidence Bound Algorithm UCB;
[0176] S0221, Define the sample mean and the upper confidence bound of each action of the radar;
[0177] S0222, The radar calculates the corresponding to each action , and selects the maximum value as the action of the current monopulse ;
[0178] S0223, The radar receives the spectrogram of the current round and obtains the return value of each monopulse from this ;
[0179] S0224, The radar counts the number of times the current action is selected , and updates the and the return of the current action through and , and the specific update formula is:
[0180]
[0181] where represents the total number of monopulses;
[0182] S023, Thompson Sampling Algorithm TS
[0183] S0231, Define the distribution mean for each action of the radar as , and the variance as
[0184] S0232, Each action of each single pulse is sampled from its respective distribution, and the radar selects the maximum value among them as the action of the current single pulse
[0185] S0233, The radar receives the spectrogram of the current single pulse and obtains the return value of the current single pulse from this
[0186] S0234, The radar updates the current action through the return to update the mean and variance, and the specific update formula is
[0187]
[0188] where is the variance of the noise
[0189] S06, Obtain the optimal radar anti-jamming cluster model
[0190] According to the multi-armed bandit method in step S05, find the result with the smallest regret value as the optimal radar anti-jamming cluster model, and determine the bandwidth and frequency for calculating the optimal radar anti-jamming cluster model based on this
[0191] S07, Simulate and calculate the optimization effect
[0192] Simulate and calculate the anti-jamming capabilities of different multi-armed bandit methods, and compare the anti-jamming capabilities of different methods through the decline rate of the regret value, the rise rate of the return value, and the final results
[0193] Simulate and calculate the optimization effect separately according to the three interference modes of the jammer
[0194] S031, Adaptive interference: The jammer makes an adaptive interference decision according to the frequency points of the radar's first 10 steps
[0195] S032, Timely forwarding: The jammer detects the frequency point of the first sub-pulse of the radar's current single pulse and chooses to continuously interfere with this frequency point
[0196] S033, Fixed interference: The jammer takes actions according to the pre-designed action probability distribution
[0197] Regarding the anti-interference ability:
[0198] 1) It is measured by the descending speed of the curve in the average regret paper image. The faster the descending speed, the stronger the anti-interference ability;
[0199] 2) It is measured by the ascending speed and the final value of the curve in the return value image. The faster the ascending speed and the larger the final return value, the stronger the anti-interference ability.)
[0200] Thus, a method for a frequency-agile radar to resist multiple interferences based on the Bandit Algorithm is obtained. Aiming at the dynamic and non-stationary electromagnetic spectrum, a frequency-agile radar is combined with the Bandit Algorithm, which has been discussed for a long time in the field of reinforcement learning, to effectively avoid the interference coverage frequency band, autonomously select the transmitting pulse carrier frequency, and thus adaptively adjust the transmitting signal frequency point of the frequency-agile radar, with the ability to adaptively resist main lobe interference and better cope with anti-main lobe interference; it improves the radar's ability to resist multiple interferences in a multi-interference environment and makes up for the problem of the weak anti-interference ability of the radar when facing a flexible and changeable interference cluster.
[0201] The present invention is not limited to the above optional implementation manners. Any person can obtain other various forms of products under the inspiration of the present invention. However, no matter what changes are made in its shape or structure, as long as the technical solutions fall within the scope defined by the claims of the present invention, they are all within the protection scope of the present invention.
Claims
1. A frequency agile radar anti-multi-interference method based on the Bandit Algorithm, characterized in that: It includes the following steps: S01. Collect the initialization system parameters of the radar Collect the parameters of the initialization system during the interaction between the radar and the interference cluster, including the number of sub-pulses of the radar, the radar bandwidth, the upper limits of the transmitting powers of the radar and the jammer, and the frequencies of the radar and the jammer; S02. Construct a radar system model Construct a model of a frequency-agile radar anti-jamming cluster based on intra-pulse frequency hopping according to the bandwidth, frequency, and power collected in step S01; and obtain the received signal after the online interaction between the frequency-agile radar anti-jamming cluster model and the interference cluster; S03. Calculate the return value and regret value of the frequency-agile radar anti-jamming cluster model Calculate the signal-to-interference-plus-noise ratio as the return value of this interaction based on the received signal of the frequency-agile radar anti-jamming cluster model in step S02; and obtain the total regret value through the return value of each pulse round; S04. Establish an objective function Combine the return value and regret value obtained in step S03, and the radar anti-jamming cluster model in step S02 to determine the objective function of the radar anti-jamming cluster model; S05. Introduce multiple multi-armed bandit methods to solve the objective function Introduce multiple multi-armed bandit methods for the objective function to optimize the objective function of the frequency-agile radar anti-jamming cluster model; S06. Obtain the optimal radar anti-jamming cluster model According to the multi-armed bandit method in step S05, find the result with the smallest regret value as the optimal radar anti-jamming cluster model, and calculate the bandwidth and frequency of the optimal radar anti-jamming cluster model; S07. Simulate and calculate the optimization effect Simulate and calculate the anti-jamming capabilities of different multi-armed bandit methods, and compare the anti-jamming capabilities of different methods through the decrease rate of the regret value, the increase rate of the return value, and the final result.
2. The frequency-agile radar anti-multi-interference method based on the Bandit Algorithm according to claim 1, wherein: The content included in step S02 is as follows: Create a frequency-agile radar anti-jamming cluster model according to the single-pulse carrier during the interaction between the radar and the interference cluster in each pulse period.
3. The frequency-agile radar anti-multi-interference method based on the Bandit Algorithm according to claim 2, characterized in that: The content included in step S02 is as follows: Create a frequency-agile radar anti-jamming cluster model according to the corresponding frequencies of multiple sub-pulses in each single pulse.
4. The frequency-agile radar anti-multi-interference method based on the Bandit Algorithm according to claim 3, wherein: The content included in step S02 is as follows: The frequency-agile radar anti-jamming cluster model includes a transmitting and receiving model of the radar and a transmitting and receiving model of the jammer.
5. The frequency agile radar anti-multi-interference method based on the Bandit Algorithm according to claim 4, characterized in that: The content included in step S02 is as follows: Create a transmitting signal model of the frequency-agile radar according to the following function formula: where K is the number of sub - pulses within a single radar pulse, a(t) is the complex envelope of the radar signal, T c is the sub - pulse duration of the radar, rect is the rectangular function, exp is the exponential function, and t is time; f k is the carrier frequency corresponding to the radar sub - pulse, f k is selected from the frequency set F = {f1, f2, …, f N}; Each radar sub - pulse randomly selects a carrier frequency from the set F to obtain the action space A of the radar R = F×F×…×F = F K ; When K = 3 and each sub-pulse has 3 frequency points, A R = 27.
6. The frequency-agile radar anti-multi-interference method based on the Bandit Algorithm according to claim 5, wherein: The content included in step S02 is as follows: Create a transmitting and receiving model of the jammer according to the following function formula: where \(v(t)\) is the complex envelope of the jammer signal, \(T\) J is the interference pulse duration, rect is the rectangular function, exp is the exponential function, \(t\) is time, \(f\) is the carrier frequency corresponding to the interference pulse, which can also be selected from the frequency set \(F=\{f_1,f_2,\ldots,f\) N \}\); Assume that the jammer does not perform intra-pulse frequency hopping; According to the three action selections of the jammer within a single pulse: spot jamming, blanket jamming, and repeater jamming; the action space A of the jammer is obtained J = F ∪ {a N+1 , a N+2}, where a N+1 , a N+2 are the actions corresponding to blanket jamming and repeater jamming, and the actions of spot jamming are all in F 7. The frequency agile radar anti-multi-interference method based on the Bandit Algorithm according to claim 6, characterized in that: The content included in step S03 is as follows: Return value: The return value of each monopulse radar is the sum of the return values of multiple sub-pulses in each monopulse; define the return value of each monopulse radar as r t ; Regret value: Calculate the regret value according to the following function formula: where r t (a) represents the reward for choosing action a, r t (a t ) represents the reward for choosing action a t ; R t Indicates the gap between the cumulative return value of the optimal action after T rounds and the cumulative return value of the actual action in T rounds.
8. The frequency-agile radar anti-multi-interference method based on the Bandit Algorithm according to claim 7, characterized in that: The return value is calculated according to the common return function: detection probability and / or signal-to-interference ratio.
9. The frequency-agile radar anti-multi-interference method based on the Bandit Algorithm according to claim 8, characterized in that: In step S05, the content included is as follows: Introduce three multi-armed bandit methods to optimize the frequency-agile radar anti-jamming cluster model; S021. Exploration-Exploitation Algorithm Exp3 based on exponential weighting; S0211, the radar selects the action a of the radar according to the current action probability distribution p t t ; S0212, the radar receives the spectrogram of the current monopulse and obtains the return value r of the current monopulse therefrom t ; S0213, calculate the cumulative return value of each action i of the radar: S0214, the radar updates the action probability distribution p by accumulating the return value S t Update the action probability distribution p t , and the specific update formula is where K represents that the radar has a total of K actions; S022. Upper Confidence Bound Algorithm UCB; S0221, define the sampling mean s of each action of the radar ti and the upper confidence limit y ti ; S0222, the radar calculates the x corresponding to each action i ti + ti , and selects the maximum value as the action a of the current monopulse t ; S0223, the radar receives the spectrogram of the current round and obtains the return value r of each single pulse from this t ; S0224, the radar counts the number of selections n of the current action i ti , and through n ti and the reward r t update the x of the current action t and y t , and the specific update formula is: where T represents the total number of single pulses; S023. Thompson Sampling Algorithm TS; S0231, define the distribution mean of each action i of the radar as μ ti , and the variance is For S0232, each action i of each single pulse is sampled from its respective distribution, and the radar selects the maximum value among them as the action a of the current single pulse. t ; S0233, the radar receives the spectrogram of the current monopulse and thereby obtains the return value r of the current monopulse t ; S0234, the radar returns r t Update the mean and variance of the current action i. The specific update formula is wherein is the variance of the noise.
10. The frequency-agile radar anti-multi-interference method based on the Bandit Algorithm according to claim 9, characterized in that: The content included in step S07 is as follows: The optimization effects are simulated and calculated separately according to the three jamming modes of the jammer; S031, Adaptive jamming: The jammer makes an adaptive jamming decision based on the frequency points in the first 10 steps of the radar; S032, Timely forwarding: The jammer detects the frequency point of the first sub-pulse of the current monopulse of the radar and chooses to jam this frequency point continuously; S033, Fixed jamming: The jammer takes actions according to the pre-designed action probability distribution.
Citation Information
Patent Citations
Ground penetrating radar intelligent inversion method based on deep learning
CN111781576A
MAB model-based FAR anti-active suppression interference strategy generation method
CN115586496A