An intelligent radar online anti-jamming method based on MAB

By adopting an intelligent anti-interference method based on MAB in radar, and using UCB decision makers to adjust the transmission strategy and process signals, the problem of difficulty in dealing with various unknown interferences in the existing technology is solved, and a more efficient and intelligent anti-interference effect is achieved.

CN119936807BActive Publication Date: 2025-06-17CHINA UNIV OF PETROLEUM (EAST CHINA) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510423086.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-17
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing online learning cognitive anti-interference method is difficult to cope with multiple unknown types of interference, and radars are separated from each other in the transmission, reception and processing links, limiting their intelligence and anti-interference performance.

Method used

The intelligent radar online anti-jamming method based on multiple alphabeta (MAB) is adopted. By building a UCB-based decision-maker, anti-jamming actions are selected and output, the radar launch strategy is adjusted, and the decision-maker status is updated through interference suppression processing and effect evaluation, and the anti-jamming strategy is optimized.

Benefits of technology

It realizes effective learning and response to many unknown types of interference, improves the anti-interference ability and intelligence of the radar, and has good robustness and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119936807B_ABST
    Figure CN119936807B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent radar online anti-jamming method based on MAB, which relates to the field of radar anti-jamming and specifically includes: parameter initialization; the decision maker selects and outputs anti-jamming actions; the radar emits signals according to the adjusted transmission strategy for target detection, and after receiving the echo signal, the selected anti-jamming action performs interference suppression processing on the signal; the processed signal result is transmitted to the anti-jamming effect evaluator to obtain the anti-jamming benefit score; the updater updates and optimizes the state of the decision maker according to the anti-jamming benefit score, and then enters the next confrontation process; if the updater detects that the jammer changes the jamming state, the current jamming state and the corresponding anti-jamming decision algorithm state parameters are stored in the jamming state memory bank, otherwise the confrontation iteration count is updated and the next confrontation is entered. The technical solution of the present invention overcomes the problem in the prior art that it is impossible to learn multiple unknown types of interference and thus impossible to perform online anti-jamming.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of radar anti - jamming, and particularly to an intelligent radar online anti - jamming method based on MAB. Background Art

[0002] The confrontation scenario between a cognitive radar and an intelligent jammer in interaction and game is as follows: on the one hand, the jammer intercepts the radar - emitted signal, extracts its key feature information, and designs a jamming strategy based on a specific decision - making method to jam the radar; on the other hand, the cognitive radar receives the jammed echo signal, conducts back - end signal processing and jamming information extraction, and adaptively adjusts subsequent transmission and reception strategies by means of intelligent anti - jamming decision - making means, so as to achieve cognitive anti - jamming detection of the radar. In this game confrontation, the key for the cognitive radar to win lies in the design of the anti - jamming decision - making method.

[0003] The core of anti - jamming decision - making is: under the condition that the jamming strategy of the jammer is unknown, it can detect the change of the jamming state, and through adaptively adjusting the radar transmission and reception strategies under each jamming state, quickly optimize and solve the optimal anti - jamming strategy under the current jamming state, so as to make the radar side occupy a longer - time and more significant confrontation advantage during the confrontation. The research on detecting the change of the jamming state has been relatively mature, so the present invention focuses on how to solve the optimal anti - jamming strategy. The current online learning cognitive anti - jamming methods face two problems: First, the existing methods mainly focus on cognitive transmission technology, and for a specific combination of transmission technology and signal processing technology in the actual environment, it is difficult to comprehensively and effectively counter multiple types of unknown jammings and it is also difficult to have good robustness in multiple scenarios. Second, the radar is fragmented in the links of cognitive transmission, cognitive reception and intelligent processing, which limits the intelligence level of the cognitive radar and cannot fully exert its anti - jamming efficiency.

[0004] Therefore, there is a need for a radar anti - jamming method that can cope with multiple unknown types of jammings and can perform online cognition. Summary of the Invention

[0005] The main purpose of the present invention is to provide an intelligent radar online anti - jamming method based on MAB to solve the problem in the prior art that it cannot learn multiple unknown types of jammings and thus cannot perform online anti - jamming.

[0006] To achieve the above object, the present invention provides an intelligent radar online anti - jamming method based on MAB, which specifically includes the following steps:

[0007] S1, the radar and the jammer start to confront, and initialize the state parameters of the anti - jamming decision - making algorithm according to the current jamming state.

[0008] S2. Construct a UCB-based decision maker that selects and outputs anti-jamming actions for adjusting the radar's transmission strategy.

[0009] S3. The radar transmits signals for target detection according to the adjusted transmission strategy. After receiving the echo signals, the selected anti-jamming actions perform interference suppression processing on the signals.

[0010] S4. The processed signal results are transmitted to the anti-jamming effect evaluator to evaluate the interference suppression effect and obtain the anti-jamming benefit score.

[0011] S5. The UCB-based updater updates and optimizes the state of the decision maker according to the anti-jamming benefit score, and then enters the next confrontation process.

[0012] S6. If the updater detects that the jammer changes the jamming state, the current jamming state and the corresponding anti-jamming decision algorithm state parameters are stored in the jamming state memory bank; otherwise, the confrontation iteration count is updated and the next confrontation is entered.

[0013] Furthermore, the anti-jamming decision algorithm state parameters in step S1 include: the mean anti-jamming action benefit, the UCB value, the action selection count, and the confrontation iteration count.

[0014] Furthermore, step S2 specifically includes the following steps:

[0015] S2.1. Scale the scale of each anti-jamming action benefit mean to the interval.

[0016] S2.2. Calculate the UCB value, and select the action corresponding to the largest UCB value as the anti-jamming action output by the decision maker:

[0017] ;

[0018] ;

[0019] ;

[0020] where, is the mean anti-jamming action benefit, is the selection count of action ; is the benefit of the -th selection of action ; is the confrontation iteration count, is the UCB value of action ; is an adjustable parameter for measuring uncertainty, is a non-zero value used to prevent the denominator from being zero. is an anti-interference action, is the maximum value function.

[0021] Furthermore, step S4 specifically includes the following steps:

[0022] S4.1, taking the number of target detections obtained after pulse compression and CFAR detection of the echo signal after interference suppression as an evaluation factor for the interference suppression effect :

[0023] .

[0024] S4.2, judging the authenticity of the identified targets, and the calculation of the evaluation factor is as follows:

[0025] ;

[0026] ;

[0027] wherein, is the impulse response width of the measured target, is the theoretical impulse response width of the target, is the speed of light, is the instantaneous operating bandwidth of the radar, is the main lobe width error.

[0028] S4.3, analyzing the pulse compression result after interference suppression using the peak sidelobe ratio PSR:

[0029] ;

[0030] wherein, is the maximum peak power in the pulse compression result, is the power of the second largest peak.

[0031] S4.4, integrating the three evaluation factors 、 and , to obtain the anti-interference benefit score :

[0032] ;

[0033] wherein, is the echo signal after being processed by the anti-interference action.

[0034] Furthermore, step S5 is specifically:

[0035] When obtaining the anti-interference benefit evaluation score After that, update the action step by step The mean value of the anti-jamming action benefit and the number of action selections of

[0036] ;

[0037] ;

[0038] Among them, is The mean value of the anti-jamming action benefit at time is The anti-jamming benefit score at time is The action at time The number of selections.

[0039] Furthermore, step S6 is specifically as follows: If the updater detects that the jammer changes the jamming state, then store the current jamming state And the corresponding anti-jamming decision algorithm state parameters in the form of parameter pairs into the jamming state memory bank; in a new round of confrontation, if the jamming state that appears matches the jamming state already existing in the jamming state memory bank, then use the anti-jamming decision algorithm state parameters corresponding to the jamming state as the initialization parameters and start a new round of confrontation; if they do not match, re-initialize according to the current anti-jamming state.

[0040] The present invention has the following beneficial effects:

[0041] The present invention comprehensively considers various anti-jamming algorithms and establishes a complete anti-jamming action library as an alternative for the interference confrontation algorithm. This anti-jamming action library covers anti-jamming actions for dealing with various types of interference. These actions can be either active, that is, adjusting the radar transmission strategy and using the conventional method for signal processing; or passive, that is, keeping the active transmission strategy unchanged and adjusting the signal processing module; or even composite, that is, combining the active transmission strategy with intelligent processing. This anti-jamming action library has scalability and replaceability and has good practical operability in engineering applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0043] Figure 1 Shows the flowchart of an intelligent radar online anti-jamming method based on MAB of the present invention.

[0044] Figure 2 Shows the execution time graph of anti-interference using the method provided by the present invention.

[0045] Figure 3 Shows the average revenue graph of the multi-interference type scenario algorithm.

[0046] Figure 4 Shows the comparison graph of the anti-interference return of CAMI-UCB and other MAB performance algorithms under noise amplitude modulation interference.

[0047] Figure 5 Shows the comparison graph of CAMI-UCB and other MAB performance algorithms under noise phase modulation interference.

[0048] Figure 6 Shows the comparison graph of CAMI-UCB and other MAB performance algorithms under noise frequency aiming interference.

[0049] Figure 7 Shows the comparison graph of CAMI-UCB and other MAB performance algorithms under intermittent sampling direct forwarding interference ISDJ.

[0050] Figure 8 Shows the comparison graph of CAMI-UCB and other MAB performance algorithms under intermittent sampling and retransmission interference ISRJ.

[0051] Figure 9 Shows the comparison graph of CAMI-UCB and other MAB performance algorithms under range deception interference. Detailed implementation manners

[0052] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0053] Embodiment 1

[0054] As Figure 1 shown, a kind of intelligent radar online anti-interference method based on MAB specifically includes the following steps:

[0055] S1. The radar and the jammer start to confront, and initialize the state parameters of the anti-interference decision algorithm according to the current interference state.

[0056] S2. Construct a UCB-based decision maker, and the decision maker selects and outputs an anti-interference action, and the anti-interference action is used to adjust the transmission strategy of the radar.

[0057] S3, the radar transmits signals according to the adjusted transmission strategy for target detection. After receiving the echo signal, the selected anti-interference action performs interference suppression processing on the signal.

[0058] S4, the processed signal result is transmitted to the anti-interference effect evaluator to evaluate the interference suppression effect and obtain the anti-interference benefit score.

[0059] S5, the UCB-based updater updates and optimizes the state of the decision maker according to the anti-interference benefit score, and then enters the next adversarial process.

[0060] S6, if the updater detects that the jammer changes the interference state, the current interference state and the corresponding anti-interference decision algorithm state parameter Key-Value are stored in the interference state memory, otherwise the number of confrontation iterations is updated and the next confrontation is entered.

[0061] Due to the uncertainty of the electromagnetic environment and the dynamic changes of the jammer's interference parameters, each anti-interference action in the anti-interference action library can be regarded as a device with a specific benefit distribution. The confrontation decision-making process based on the complete anti-interference action library can be modeled as a reinforcement learning model MAB. The anti-interference action library contains anti-interference algorithms such as adaptive sidelobe cancellation and narrow pulse rejection. In the case of fuzzy cognition of jammer interference, the radar makes decisions to select the optimal anti-interference action based on the anti-interference benefit history of each action.

[0062] The method proposed in this invention focuses on the problem of robust target detection within the radar coherent processing interval (CPI). In each CPI frame, the radar anti-interference decision is optimized in real time, and each CPI frame outputs an anti-interference action, which will act on the next frame. Through iterative calculation, the optimal anti-interference action is solved to pursue the highest possible anti-interference benefit return.

[0063] From the perspective of the time axis, the execution process of the method provided by the present invention is as follows: Figure 2 As shown. The green area indicates the time that the radar can be used for anti-interference processing, and the green part is further divided into two stages: the first stage is used to select anti-interference actions and adjust the radar transmission strategy accordingly; the second stage is used to perform interference suppression processing on the received signal and subsequent related operations. The blue area represents the time for the radar to perform routine operations such as transmission and reception. t represents the duration of a CPI frame, that is, the time period for one radar to confront the jammer. This anti-interference decision framework has a wide range of applicability, is compatible with passive and active interference suppression algorithms, and connects cognitive transmission and intelligent processing modules.

[0064] Specifically, the state parameters of the anti-interference decision algorithm in step S1 include: the mean anti-interference action benefit, the UCB value, the number of action selections and the number of adversarial iterations.

[0065] Specifically, in the decision-making process, the UCB algorithm always selects the action with the highest estimated reward each time. Under this mechanism, for the overestimation of action rewards, it will either be quickly corrected because the action is frequently selected, or even if overestimated, it will not cause serious consequences (because it will not be selected); however, the underestimation of action rewards is extremely dangerous and may lead to a large regret value. Based on this, the UCB series of algorithms introduce the upper confidence bound to measure the uncertainty of action rewards. Step S2 specifically includes the following steps:

[0066] S2.1, scale the mean of each anti-jamming action reward to the interval.

[0067] S2.2, calculate the UCB value, and select the action corresponding to the largest UCB value as the anti-jamming action output by the decision maker:

[0068] ;

[0069] ;

[0070] ;

[0071] where is the mean of the anti-jamming action rewards, is the number of times the action is selected, is the reward of the th selection of the action , is the number of anti-jamming iterations, is the UCB value of the action , is an adjustable parameter for measuring uncertainty, is a non-zero value used to prevent the denominator from being zero, is the anti-jamming action, is the maximum value function. The set of anti-jamming actions constitutes the anti-jamming action library.

[0072] Specifically, for the scenario of one radar, one jammer, and one target, the evaluation indicators of the evaluator designed by the present invention examine three key factors, namely: the number of detected targets, the impulse response width (IRW), and the peak sidelobe ratio (PSR). Step S4 specifically includes the following steps:

[0073] S4.1, the number of detected targets obtained after pulse compression and constant false alarm rate (CFAR) detection of the echo signal after interference suppression As an evaluation factor for the interference suppression effect :

[0074] 。

[0075] S4.2. For radar target detection, its IRW is determined by the bandwidth of the transmitted signal, that is . Deceptive interference will introduce false targets, and these false targets may exceed the real targets. However, the characteristics of the jammer determine that the bandwidth of the interference component is much smaller than the bandwidth of the real detection signal. Therefore, the IRW formed by the false targets will be broadened relative to the IRW of the real targets. According to this characteristic, the authenticity of the identified targets can be judged. To judge the authenticity of the identified targets, the evaluation factor is calculated as follows:

[0076] ;

[0077] ;

[0078] Wherein, is the impulse response width of the measured target, is the theoretical impulse response width of the target, is the speed of light, is the instantaneous operating bandwidth of the radar, is the main lobe width error.

[0079] S4.3. Use the PSR factor to analyze the pulse compression result after interference suppression. PSR is the ratio of the maximum peak power to the second maximum peak power of the identified target, and it can simulate the signal-to-jamming ratio (SJR) of the pulse compression result under the condition of no prior interference. This is the most intuitive manifestation of the interference suppression effect. Especially when facing deceptive interference, when the transmit power and the target RCS characteristics do not change, the larger the SJR, the smaller the probability of detecting false targets and the smaller the impact on radar target detection. Use the peak sidelobe ratio PSR to analyze the pulse compression result after interference suppression:

[0080] ;

[0081] Wherein, is the maximum peak power in the pulse compression result, is the power of the second largest peak.

[0082] S4.4. Integrate the three evaluation factors 、 and to obtain the anti-interference benefit score :

[0083] ;

[0084] wherein, is the echo signal after anti-interference action processing. Under the revenue function , the selection of non-optimal actions will bring a large regret value to the radar, and when the initial action revenue is poor, encouraging the radar to adjust the action in a timely manner helps to improve the rapid revenue acquisition and convergence ability of the actual online anti-interference decision-making.

[0085] Specifically, step S5 is specifically as follows:

[0086] After obtaining the anti-interference revenue evaluation score , update the anti-interference action revenue mean value and the action selection times of action in a step-by-step manner:

[0087] ;

[0088] ;

[0089] wherein, is the anti-interference action revenue mean value at time , is the anti-interference revenue score at time , is the selection times of action at time

[0090] Specifically, step S6 is specifically as follows: If the updater detects that the jammer changes the interference state, store the current interference state and the corresponding anti-interference decision algorithm state parameters in the interference state memory bank in the form of parameter pairs; in a new round of confrontation, if the interference state that appears matches the interference state existing in the interference state memory bank, use the anti-interference decision algorithm state parameters corresponding to the interference state as the initialization parameters and start a new round of confrontation; if they do not match, re-initialize according to the current anti-interference state.

[0091] To verify the beneficial effects of the method of the present invention, the following simulation experiments are carried out:

[0092] The simulation experiment sets up a scenario that includes a fire control radar, a target, and a jammer. The radar emits a Linear Frequency Modulation (LFM) signal. The jammer can implement sidelobe or mainlobe jamming and has 6 different jamming states, covering a variety of typical jamming types: blocking noise amplitude modulation jamming, blocking noise phase modulation jamming, narrowband frequency tracking noise phase modulation jamming (hereinafter referred to as noise amplitude modulation, noise phase modulation, and noise frequency tracking jamming respectively), Interrupted-sampling and Direct Repeater Jamming (ISDJ), Interrupted-sampling Repeater Jamming (ISRJ), and range deception jamming. The specific parameter settings of the radar, target, and jammer are shown in Table 1. When the jammer performs range deception jamming, its pulse width, bandwidth, and center frequency are consistent with the radar emission parameters.

[0093] Five passive anti-jamming methods are simulated and implemented in the anti-jamming action library, including Sidelobe Cancellation (SLC), coherent integration, noise interference reconstruction cancellation, narrow pulse reconstruction cancellation, and narrow pulse rejection; at the same time, two active anti-jamming methods are implemented, including pulse cover and random frequency agility, for a total of 7 anti-jamming actions. Each anti-jamming action contains adjustable parameters, such as the number of sub-antennas in SLC, the narrow pulse width recognition threshold in narrow pulse rejection, etc.

[0094] To evaluate the performance of each anti-jamming action, each anti-jamming action was repeatedly tested 200 times under different jamming states, and the mean values of their anti-jamming benefit scores are as shown. The optimal results in the current experiment are marked in bold in Table 2, and the sub-optimal results in the current experiment are underlined; S.L.I. (Side Lobe Interference) represents the anti-jamming effect of SLC when the jammer performs sidelobe jamming.

[0095] Table 1 Radar, Target, and Jammer Parameter Settings

[0096]

[0097] Table 2 Average Benefits of Anti-Jamming Actions

[0098]

[0099] The anti-interference method proposed in this invention is named CAMI-UCB (Cognitive Algorithm for Multi-type Interferences based on UCB). To verify its effectiveness against consecutive multiple interference states, the jammer successively uses noise frequency tracking, range deception, noise frequency tracking, and ISRJ to interfere with the radar. The number of continuous iterations for each interference state is 200, 100, 200, and 100 respectively, with a total of 600 confrontation iterations. The experiment compares three online learning anti-interference methods, namely: CAMI-UCB proposed in this invention, the frequency agility online decision-making algorithm RAFA-EXP3++ based on MAB, and the multi-domain adaptive active emission algorithm TSDT. In the experiment, TSDT can optimize the parameters in two domains, namely frequency and pulse width. Each algorithm conducts 100 repeated experiments with the jammer, and the experimental results are as Figure 3 shown. The horizontal axis is the number of iterative confrontations, and the vertical axis is the average gain of the repeated experiments. The experimental results show that the CAMI-UCB algorithm can quickly converge to an effective anti-interference action when facing a new interference state. For example, when first dealing with noise frequency tracking interference (Epoch = 1~200), range deception (Epoch = 200~300), and ISRJ (Epoch = 500~600), the algorithm can converge within about 50 steps; when facing an interference state that has been encountered (Epoch = 300~500), it can quickly restore the corresponding algorithm state and avoid repeated solutions. Compared with RAFA-EXP3++ and TSTD, CAMI-UCB has the following advantages:

[0100] Quickly solving high-gain actions: When there are actions with greater anti-interference potential in the action library, CAMI-UCB can quickly converge to a more stable anti-interference action with a higher average gain. For example, in the experiment when dealing with noise frequency tracking interference, CAMI-UCB can stably converge to a higher-gain action (noise reconstruction cancellation, coherent accumulation).

[0101] Adaptability to multi-type interferences: CAMI-UCB can effectively deal with multi-type interference states, solving the problem that traditional radars cannot cope with non-targeted interferences due to relying on specific anti-interference algorithms. For example, when dealing with ISRJ interference, both RAFA-EXP3++ and TSTD, the two active emission anti-interference methods, lose their anti-interference capabilities, while CAMI-UCB can solve effective anti-interference actions against ISRJ from the action library.

[0102] In addition, the RAFA-EXP3++ and TSTD algorithms can also be used as alternative algorithms in the anti-interference action library. By combining with the joint intelligent processing technology for anti-ISRJ, the radar strategy can avoid the ISRJ interference band and further eliminate the interference.

[0103] To verify the performance advantage of CAMI-UCB over other classical MAB algorithms, the present invention conducted experimental tests under six different interference states set by the jammer, and compared the anti-jamming performance of the upper confidence bound algorithm

[0104] UCB, Thompson sampling algorithm TS, decaying epsilon-greedy algorithm DEG, and the multi-type cognitive anti-jamming method CAMI. Among them, TS assumes that the reward distribution of the device follows a Gaussian distribution, and the sampling distribution of the th action at the th iteration is:

[0105] ;

[0106] where represents the Gaussian distribution, is an adjustable parameter; in the experiment, is taken.

[0107] In DEG, the decay function is set to:

[0108] ;

[0109] where is the th iteration of value, is the initial value, set to 0.6. Each algorithm conducts 200 iterations of confrontation with the jammer under each interference state and repeats the experiment 10 times. Its average reward (cumulative gain) is as Figures 4 - 9 shown. The experimental results show that UCB performs better than TS and DEG in most cases. Although TS and UCB have similar performance, since UCB does not require strong assumptions about the reward distribution of anti-jamming actions, it has better adaptability and superior performance.

[0110] Aiming at the problems of interference from multi-interference type jammers, complex battlefield electromagnetic environment, and lack of prior knowledge, the present invention constructs a complete anti-jamming action library, models the reward distribution of each anti-jamming action as a device, and thus transforms the confrontation problem based on the anti-jamming action library into an MAB problem. Based on the working characteristics of the jammer, an online learning anti-jamming framework CAMI is proposed, and an anti-jamming reward evaluation method applicable to single radar, single multi-interference type jammer, and single target scenario is designed. The algorithm uses the upper confidence bound (UCB) algorithm to solve the exploration and exploitation trade-off problem.

[0111] In the simulation experiment, the present invention implemented a jammer model including six types of jamming, namely, blocking noise amplitude modulation jamming, blocking noise phase modulation jamming, frequency-seeking noise phase modulation jamming, intermittent sampling direct forwarding jamming (ISDJ), intermittent sampling relaying jamming (ISRJ), and range deception jamming, and verified the effectiveness of CAMI-UCB in continuous multi-type jamming scenarios. The experimental results show that even when a single anti-jamming method fails, CAMI-UCB can still effectively counteract jamming and quickly solve the anti-jamming action that best suits the current jamming state. In the comparative tests with other MAB-based algorithms, UCB demonstrated better performance and adaptability than Decaying Epsilon Greedy (DEG) and Thompson Sampling (TS). In summary, CAMI-UCB has the following advantages:

[0112] Combined active and passive anti-jamming technology: By combining active and passive anti-jamming technologies, the synergistic effect of radar transmission technology and signal processing technology is fully utilized to obtain higher anti-jamming benefits.

[0113] Adaptability to multiple jamming types: It can effectively cope with multiple jamming types and optimize and solve the optimal anti-jamming action at a low computational cost and high speed.

[0114] Embodiment 2

[0115] The present invention also provides an intelligent radar online anti-jamming system based on MAB, which includes: a decision maker, an evaluator, and an updater connected in sequence; the decision maker selects and outputs an anti-jamming action, which is used to adjust the radar's transmission strategy, and the anti-jamming action is stored in the anti-jamming action library; the radar transmits signals for target detection according to the adjusted transmission strategy, and after receiving the echo signal, the selected anti-jamming action performs interference suppression processing on the signal; the processed signal result is transmitted to the evaluator, and the evaluator gives an anti-jamming benefit score; the updater updates and optimizes the state of the decision maker according to the anti-jamming benefit score, and then enters the next confrontation process. If the updater detects that the jammer changes the jamming state, the current jamming state E and the state parameters of the anti-jamming decision algorithm are stored in the jamming state memory library in the form of a Key-Value pair. In a new round of confrontation, if the jamming state E that appears matches the Key in the jamming state memory library, the corresponding Value is used to initialize the system and start a new round of confrontation; if not, it is re-initialized according to the current jamming state E.

[0116] Certainly, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.

Claims

1. An online anti-interference method for intelligent radar based on MAB, characterized in that: The specific steps include: S1, the radar and the jammer begin to confront each other, and the state parameters of the counter-interference decision algorithm are initialized according to the current interference state E; S2, build a UCB-based decision maker, which selects and outputs anti-interference actions, which are used to adjust the radar's transmission strategy; S3, the radar transmits signals according to the adjusted transmission strategy for target detection, and after receiving the echo signal, the selected anti-interference action performs interference suppression processing on the signal; S4, the processed signal result is transmitted to the anti-interference effect evaluator to evaluate the interference suppression effect and obtain the anti-interference benefit score; S5, the UCB-based updater updates and optimizes the state of the decision maker according to the anti-interference benefit score, and then enters the next confrontation process; S6, if the updater detects that the jammer changes the jamming state, the current jamming state and the corresponding state parameters of the anti-jamming decision algorithm are stored in the jamming state memory bank, otherwise the number of confrontation iterations is updated and the next confrontation is entered; Step S4 specifically includes the following steps: S4.1, the number of target detections obtained after pulse compression and CFAR detection of the echo signal after interference suppression As an evaluation factor for interference suppression effect : ; S4.2, judge the authenticity of the identified target and evaluate the factors The calculation formula is as follows: ; ; in, To measure the impulse response width of the target, is the target theoretical impulse response width, is the speed of light, is the instantaneous working bandwidth of the radar, is the main lobe width error; S4.3, use the peak sidelobe ratio PSR to analyze the pulse compression results after interference suppression: ; in, is the maximum peak power in the pulse compression result, is the power of the next largest peak; S4.4, integration of three evaluation factors , and , and get the anti-interference benefit score : ; in, It is the echo signal after anti-interference action processing.

2. According to claim 1, a MAB-based intelligent radar online anti-interference method is characterized in that: The state parameters of the anti-interference decision algorithm in step S1 include: the mean anti-interference action benefit, the UCB value, the number of action selections and the number of adversarial iterations.

3. The MAB-based intelligent radar online anti-interference method according to claim 1 is characterized in that: Step S2 specifically includes the following steps: S2.1, scale the mean value of each anti-interference action benefit to interval; S2.2, calculate the UCB value, and select the action corresponding to the one with the largest UCB value as the anti-interference action output by the decision maker: ; ; ; in, is the mean benefit of the anti-interference action, It's action The number of selections, It's action No. The profit of the selected is the number of adversarial iterations, It's action The UCB value, is an adjustable parameter to measure uncertainty, Is a non-zero value, used to prevent the denominator from being 0. It is an anti-interference action. is the maximum value function.

4. The MAB-based intelligent radar online anti-interference method according to claim 1 is characterized in that: Step S5 is specifically as follows: Obtaining the anti-interference benefit evaluation score Then, update the action in a step-by-step manner The mean anti-interference action benefit and the number of action selections: ; ; in, for The mean value of the anti-interference action benefit at the moment, for The anti-interference benefit score at the moment, for Moment of action The number of selections.

5. The MAB-based intelligent radar online anti-interference method according to claim 1 is characterized in that: Step S6 is as follows: if the updater detects that the jammer has changed its jamming state, the current jamming state is updated. and the corresponding anti-interference decision algorithm state parameters are stored in the interference state memory bank in the form of parameter pairs; in a new round of confrontation, if the interference state that appears matches the existing interference state in the interference state memory bank, the anti-interference decision algorithm state parameters corresponding to the interference state are used as initialization parameters, and a new round of confrontation is started; if they do not match, they are reinitialized according to the current anti-interference state.