Intelligent radar online anti-interference method based on MAB
By constructing an intelligent radar anti-jamming system based on the reinforcement learning method with the upper confidence boundary of multiple arms, the response problems of various unknown interference is solved, the radar's anti-jamming ability and intelligence are improved, and the optimization strategy solution for multiple scenarios is realized.
Patent Information
- Application Number
- CN202510423086.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The prior art is difficult to deal with many unknown types of interference, and the intelligence of cognitive radar is limited, so it is unable to fully exert anti-interference performance.
Using a reinforcement learning method based on multi-arm with upper confidence boundary (UCB), an intelligent radar anti-jamming system is built. Through the combination of decision makers, evaluators and updaters, radar transmission and reception strategies are optimized, and a complete anti-jamming action library is established, covering a variety of interference types of anti-jamming actions.
It realizes effective response to multiple types of interference, improves the anti-interference ability of the radar, has good robustness and rapid response capabilities, and can optimize and solve the optimal anti-interference strategy in multiple scenarios.
Smart Images

Figure CN119936807A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of radar anti-interference, and in particular to an online anti-interference method for an intelligent radar based on MAB. Background Art
[0002] The interactive game between cognitive radar and intelligent jammer is as follows: on the one hand, the jammer detects the radar transmission signal, extracts its key feature information, and designs jamming strategies based on specific decision-making methods to jam the radar; on the other hand, the cognitive radar receives the jammed echo signal, performs back-end signal processing and jamming information extraction, and uses intelligent anti-interference decision-making methods to adaptively adjust subsequent transmission and reception strategies, thereby realizing the radar's cognitive anti-interference detection. In this game confrontation, the key to the victory of cognitive radar lies in the design of anti-interference decision-making methods.
[0003] The core of anti-interference decision-making is: under the condition that the jammer's interference strategy is unknown, it can detect the change of the interference state, and adaptively adjust the radar transmission and reception strategy under each interference state, quickly optimize and solve the optimal anti-interference strategy under the current interference state, so that the radar party can occupy a longer and more significant confrontation advantage during the confrontation. The research on detecting changes in interference states is relatively mature, so the present invention focuses on how to solve the optimal anti-interference strategy. The current online learning cognitive anti-interference method faces two problems: First, the existing methods mainly focus on cognitive transmission technology, and the specific combination of transmission technology and signal processing technology is difficult to fully and effectively confront various types of unknown interference in the actual environment, and it is also difficult to have good robustness in multiple scenarios. Second, the radar is separated from each other in the links of cognitive transmission, cognitive reception and intelligent processing, which limits the intelligence of cognitive radar and cannot give full play to its anti-interference performance.
[0004] Therefore, there is a need for a radar anti-interference method that can cope with multiple unknown types of interference and can recognize it online. Summary of the invention
[0005] The main purpose of the present invention is to provide an online anti-interference method for an intelligent radar based on MAB, so as to solve the problem that the prior art cannot learn multiple unknown types of interference and thus cannot perform online anti-interference.
[0006] To achieve the above object, the present invention provides an online anti-interference method for an intelligent radar based on MAB, which specifically comprises the following steps: S1, the radar and the jammer begin to confront each other, and the state parameters of the jamming decision algorithm are initialized according to the current jamming state.
[0007] S2, build a UCB-based decision maker, which selects and outputs anti-interference actions, which are used to adjust the radar's transmission strategy.
[0008] S3, the radar transmits signals according to the adjusted transmission strategy for target detection. After receiving the echo signal, the selected anti-interference action performs interference suppression processing on the signal.
[0009] S4, the processed signal result is transmitted to the anti-interference effect evaluator to evaluate the interference suppression effect and obtain the anti-interference benefit score.
[0010] S5, the UCB-based updater updates and optimizes the state of the decision maker according to the anti-interference benefit score, and then enters the next confrontation process;
[0011] S6, if the updater detects that the jammer changes the jamming state, the current jamming state and the corresponding state parameters of the anti-jamming decision algorithm are stored in the jamming state memory bank, otherwise the number of confrontation iterations is updated and the next confrontation is entered.
[0012] Furthermore, the state parameters of the anti-interference decision algorithm in step S1 include: the mean anti-interference action benefit, the UCB value, the number of action selections and the number of adversarial iterations.
[0013] Furthermore, step S2 specifically includes the following steps: S2.1, scale the mean value of each anti-interference action benefit to interval.
[0014] S2.2, calculate the UCB value, and select the action corresponding to the one with the largest UCB value as the anti-interference action output by the decision maker: ; ; ; in, is the mean benefit of the anti-interference action, It's action The number of selections, It's action No. The profit of the second choice, is the number of adversarial iterations, It's action The UCB value, is an adjustable parameter to measure uncertainty, Is a non-zero value, used to prevent the denominator from being 0. It is an anti-interference action. is the maximum value function.
[0015] Furthermore, step S4 specifically includes the following steps: S4.1, the number of target detections obtained after pulse compression and CFAR detection of the echo signal after interference suppression As an evaluation factor for interference suppression effect : .
[0016] S4.2, judge the authenticity of the identified target and evaluate the factors The calculation formula is as follows: ; ; in, To measure the impulse response width of the target, is the target theoretical impulse response width, is the speed of light, is the instantaneous working bandwidth of the radar, is the main lobe width error.
[0017] S4.3, use the peak sidelobe ratio PSR to analyze the pulse compression results after interference suppression: ; in, is the maximum peak power in the pulse compression result, It is the power of the next largest peak.
[0018] S4.4, integration of three evaluation factors , and , and get the anti-interference benefit score : ; in, It is the echo signal after anti-interference action processing.
[0019] Furthermore, step S5 is specifically as follows: Obtaining the anti-interference benefit evaluation score Then, update the action in a step-by-step manner The mean anti-interference action benefit and the number of action selections: ; ; in, for The mean value of the anti-interference action benefit at the moment, for The anti-interference benefit score at the moment, for Moment of action The number of selections.
[0020] Further, step S6 is specifically as follows: if the updater detects that the jammer changes the interference state, the current interference state is changed. and the corresponding anti-interference decision algorithm state parameters are stored in the interference state memory bank in the form of parameter pairs; in a new round of confrontation, if the interference state that appears matches the existing interference state in the interference state memory bank, the anti-interference decision algorithm state parameters corresponding to the interference state are used as initialization parameters, and a new round of confrontation is started; if they do not match, they are reinitialized according to the current anti-interference state.
[0021] The present invention has the following beneficial effects: The present invention comprehensively considers a variety of anti-interference algorithms and establishes a complete anti-interference action library as an alternative to the interference countermeasure algorithm. The anti-interference action library covers anti-interference actions for dealing with a variety of interference types. These actions can be active, that is, adjusting the radar transmission strategy, and using conventional methods for signal processing; they can also be passive, that is, the active transmission strategy remains unchanged and the signal processing module is adjusted; or even composite, that is, the active transmission strategy is combined with intelligent processing. This anti-interference action library is scalable and replaceable, and has good practicality in engineering applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the specific implementation of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the specific implementation or the prior art description. Obviously, the drawings described below are some implementations of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings: Figure 1 A flow chart of an online anti-interference method of an intelligent radar based on MAB of the present invention is shown.
[0023] Figure 2 The figure shows the execution time diagram of anti-interference using the method provided by the present invention.
[0024] Figure 3 The average benefit graph of the algorithm for multiple interference type scenarios is shown.
[0025] Figure 4 The anti-interference return comparison chart of CAMI-UCB and other MAB performance algorithms under noise amplitude modulation interference is shown.
[0026] Figure 5A comparison chart of CAMI-UCB and other MAB performance algorithms under noise phase modulation interference is shown.
[0027] Figure 6 A comparison chart of CAMI-UCB and other MAB performance algorithms under noise aiming interference is shown.
[0028] Figure 7 A comparison chart of CAMI-UCB and other MAB performance algorithms under intermittent sampling direct forwarding interference ISDJ is shown.
[0029] Figure 8 A comparison chart of CAMI-UCB and other MAB performance algorithms under intermittent sampling forwarding interference ISRJ is shown.
[0030] Fig. 9 A comparison chart of CAMI-UCB and other MAB performance algorithms under range deception interference is shown. DETAILED DESCRIPTION
[0031] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0032] Embodiment 1 like Figure 1 The MAB-based intelligent radar online anti-interference method shown in the figure specifically includes the following steps: S1, the radar and the jammer begin to confront each other, and the state parameters of the jamming decision algorithm are initialized according to the current jamming state.
[0033] S2, build a UCB-based decision maker, which selects and outputs anti-interference actions, which are used to adjust the radar's transmission strategy.
[0034] S3, the radar transmits signals according to the adjusted transmission strategy for target detection. After receiving the echo signal, the selected anti-interference action performs interference suppression processing on the signal.
[0035] S4, the processed signal result is transmitted to the anti-interference effect evaluator to evaluate the interference suppression effect and obtain the anti-interference benefit score.
[0036] S5, the UCB-based updater updates and optimizes the state of the decision maker according to the anti-interference benefit score, and then enters the next adversarial process.
[0037] S6, if the updater detects that the jammer changes the interference state, the current interference state and the corresponding anti-interference decision algorithm state parameter Key-Value are stored in the interference state memory, otherwise the number of confrontation iterations is updated and the next confrontation is entered.
[0038] Due to the uncertainty of the electromagnetic environment and the dynamic changes of the jammer's interference parameters, each anti-interference action in the anti-interference action library can be regarded as a device with a specific benefit distribution. The confrontation decision-making process based on the complete anti-interference action library can be modeled as a reinforcement learning model MAB. The anti-interference action library contains anti-interference algorithms such as adaptive sidelobe cancellation and narrow pulse rejection. In the case of fuzzy cognition of jammer interference, the radar makes decisions to select the optimal anti-interference action based on the anti-interference benefit history of each action.
[0039] The method proposed in this invention focuses on the problem of robust target detection within the radar coherent processing interval (CPI). In each CPI frame, the radar anti-interference decision is optimized in real time, and each CPI frame outputs an anti-interference action, which will act on the next frame. Through iterative calculation, the optimal anti-interference action is solved to pursue the highest possible anti-interference benefit return.
[0040] From the perspective of the time axis, the execution process of the method provided by the present invention is as follows: Figure 2 As shown. The green area indicates the time that the radar can be used for anti-interference processing, and the green part is further divided into two stages: the first stage is used to select anti-interference actions and adjust the radar transmission strategy accordingly; the second stage is used to perform interference suppression processing on the received signal and subsequent related operations. The blue area represents the time for the radar to perform routine operations such as transmission and reception. t represents the duration of a CPI frame, that is, the time period for one radar to confront the jammer. This anti-interference decision framework has a wide range of applicability, is compatible with passive and active interference suppression algorithms, and connects cognitive transmission and intelligent processing modules.
[0041] Specifically, the state parameters of the anti-interference decision algorithm in step S1 include: the mean anti-interference action benefit, the UCB value, the number of action selections and the number of adversarial iterations.
[0042] Specifically, in the decision-making process, the UCB algorithm will select the action with the highest estimated benefit each time. Under this mechanism, the overestimation of the action benefit will either be quickly corrected because the action is frequently selected, or even if it is overestimated, it will not have serious consequences (because it will not be selected); however, underestimation of the action benefit is extremely dangerous and may lead to a large regret value. Based on this, the UCB series of algorithms introduces an upper confidence bound to measure the uncertainty of the action benefit. Step S2 specifically includes the following steps:
[0043] S2.1, scale the mean value of each anti-interference action benefit to interval.
[0044] S2.2, calculate the UCB value, and select the action corresponding to the one with the largest UCB value as the anti-interference action output by the decision maker: ; ; ; in, is the mean benefit of the anti-interference action, It's action The number of selections, It's action No. The profit of the selected is the number of adversarial iterations, It's action The UCB value, is an adjustable parameter to measure uncertainty, Is a non-zero value, used to prevent the denominator from being 0. It is an anti-interference action. is the maximum value function. The set of anti-interference actions constitutes the anti-interference action library.
[0045] Specifically, for a scenario of a radar, a jammer, and a target, the evaluation index of the evaluator designed by the present invention examines three key factors, namely: the number of target detections, the impulse response width (IRW), and the peak sidelobe ratio (PSR). Step S4 specifically includes the following steps: S4.1, the number of target detections obtained after pulse compression and CFAR detection of the echo signal after interference suppression As an evaluation factor for interference suppression effect : .
[0046] S4.2, for radar target detection, its IRW is determined by the bandwidth of the transmitted signal, that is, Deceptive jamming will introduce false targets, which may exceed the real targets. However, the characteristics of the jammer determine that the bandwidth of the jammer component is much smaller than the bandwidth of the real detection signal. Therefore, the IRW formed by the false target will be wider than the IRW of the real target. Based on this characteristic, the authenticity of the identified target can be judged. The calculation formula is as follows: ; ; in, To measure the impulse response width of the target, is the target theoretical impulse response width, is the speed of light, is the instantaneous working bandwidth of the radar, is the main lobe width error.
[0047] S4.3, use the PSR factor to analyze the pulse compression results after interference suppression. PSR is the ratio of the maximum peak and sub-peak power of the identified target. It can simulate the signal-to-jam ratio (SJR) of the pulse compression result without interference a priori. This is the most intuitive reflection of the interference suppression effect. Especially in the face of deceptive interference, when the transmission power and the target RCS characteristics do not change, the larger the SJR, the smaller the probability of false targets being detected, and the smaller the impact on radar target detection. Use the peak sidelobe ratio PSR to analyze the pulse compression results after interference suppression: ; in, is the maximum peak power in the pulse compression result, It is the power of the next largest peak.
[0048] S4.4, integration of three evaluation factors , and , and get the anti-interference benefit score : ; in, is the echo signal after anti-interference action processing. In the benefit function In this case, the selection of non-optimal actions will bring a large regret value to the radar, and when the initial action benefit is poor, timely encouraging the radar to adjust its action will help improve the rapid benefit acquisition and convergence capabilities of the actual online anti-interference decision-making.
[0049] Specifically, step S5 is as follows: Obtaining the anti-interference benefit evaluation score Then, update the action in a step-by-step manner The mean anti-interference action benefit and the number of action selections: ; ; in, for The mean value of the anti-interference action benefit at the moment, for The anti-interference benefit score at the moment, for Moment of action The number of selections.
[0050] Specifically, step S6 is as follows: if the updater detects that the jammer changes the jamming state, the current jamming state is changed. and the corresponding anti-interference decision algorithm state parameters are stored in the interference state memory bank in the form of parameter pairs; in a new round of confrontation, if the interference state that appears matches the existing interference state in the interference state memory bank, the anti-interference decision algorithm state parameters corresponding to the interference state are used as initialization parameters, and a new round of confrontation is started; if they do not match, they are reinitialized according to the current anti-interference state.
[0051] In order to verify the beneficial effects of the method of the present invention, the following simulation experiments are carried out: The simulation experiment sets up a scenario including a fire control radar, a target and a jammer. The radar transmits a linear frequency modulation (LFM) signal. The jammer can implement sidelobe or mainlobe interference and has 6 different interference states, covering a variety of typical interference types: blocking noise amplitude modulation interference, blocking noise phase modulation interference, narrowband aiming frequency noise phase modulation interference (hereinafter referred to as noise amplitude modulation, noise phase modulation, noise aiming frequency interference), interrupted sampling and direct repeater jamming (ISDJ), interrupted sampling and forwarding jamming (ISRJ) and distance deception jamming. The specific parameter settings of the radar, target and jammer are shown in Table 1. When the jammer performs distance deception jamming, its pulse width, bandwidth and center frequency are consistent with the radar transmission parameters.
[0052] The anti-interference action library simulates and implements five passive anti-interference methods, including sidelobe cancellation (SLC), coherent accumulation, noise interference reconstruction cancellation, narrow pulse reconstruction cancellation, and narrow pulse elimination; at the same time, two active anti-interference methods are implemented, including pulse masking and random frequency agility, for a total of seven anti-interference actions. Each anti-interference action contains adjustable parameters, such as the number of sub-antennas in SLC and the narrow pulse width recognition threshold in narrow pulse elimination.
[0053] To evaluate the performance of each anti-interference action, each anti-interference action was tested 200 times under different interference conditions. The mean anti-interference benefit score is as follows: Table 2 uses bold to mark the best results in the current experiment, and underlines to mark the suboptimal results in the current experiment; SLI (Side Lobe Interference) represents the anti-interference effect of SLC when the jammer performs side lobe interference.
[0054] Table 1 Radar, target and jammer parameter settings
[0055] Table 2 Average benefits of anti-interference actions
[0056] The present invention names the proposed anti-interference method CAMI-UCB (Cognitive Algorithm for Multi-type Interferences based on UCB). In order to verify its effectiveness in countering multiple continuous interference states, the jammer uses noise aiming, distance deception, noise aiming and ISRJ to interfere with the radar in turn. The number of continuous iterations of each interference state is 200, 100, 200, and 100, respectively, with a total of 600 confrontation iterations. The experiment compares three online learning anti-interference methods, namely: CAMI-UCB proposed in the present invention, the frequency agile online decision algorithm RAFA-EXP3++ based on MAB, and the multi-domain adaptive active transmission algorithm TSDT. In the experiment, TSDT can optimize the parameters of the two domains of frequency and pulse width. Each algorithm and the jammer were repeated 100 times, and the experimental results are as follows. Figure 3 As shown in the figure, the horizontal axis is the number of iterative confrontations, and the vertical axis is the average benefit of repeated experiments. The experimental results show that the CAMI-UCB algorithm can quickly converge to an action that effectively counteracts interference when facing a new interference state. For example, when dealing with noise aiming frequency interference (Epoch=1~200), distance deception (Epoch=200~300) and ISRJ (Epoch=500~600) for the first time, the algorithm can converge within about 50 steps; when facing interference states that have been encountered (Epoch=300~500), it can quickly restore the corresponding algorithm state to avoid repeated solutions. Compared with RAFA-EXP3++ and TSTD, CAMI-UCB has the following advantages:
[0057] Quickly solve high-yield actions: When there are actions with greater anti-interference potential in the action library, CAMI-UCB can quickly converge to more stable anti-interference actions with higher average benefits. For example, in the experiment, when dealing with noise aiming frequency interference, CAMI-UCB can stably converge to actions with higher benefits (noise reconstruction cancellation, coherent accumulation).
[0058] Adaptability to multiple types of interference: CAMI-UCB can effectively deal with multiple types of interference states, solving the problem that traditional radars cannot deal with non-targeted interference due to their reliance on specific anti-interference algorithms. For example, when dealing with ISRJ interference, both RAFA-EXP3++ and TSTD active anti-interference methods lose their ability to counteract, while CAMI-UCB can solve effective anti-interference actions against ISRJ from the action library.
[0059] In addition, RAFA-EXP3++ and TSTD algorithms can also be used as alternative algorithms in the anti-interference action library. By combining with the joint anti-ISRJ intelligent processing technology, the radar strategy can avoid the ISRJ interference band and further eliminate interference.
[0060] In order to verify the performance advantage of CAMI-UCB over other classic MAB algorithms, the present invention conducted experimental tests under 6 different jammer conditions and compared the confidence interval upper bound algorithm.
[0061] The anti-interference performance of UCB, Thompson sampling algorithm TS, attenuated epsilon greedy algorithm DEG and multi-type cognitive anti-interference method CAMI. Among them, TS assumes that the income distribution of the equipment follows a Gaussian distribution. The action in The sampling distribution at the iteration is: ; in, represents a Gaussian distribution, is an adjustable parameter; .
[0062] The attenuation function in DEG is set as: ; in, For the Iteration value, For initial The value is set to 0.6. Each algorithm conducts 200 iterations of confrontation with the jammer under various jamming conditions and repeats the experiment 10 times. The average return (cumulative benefit) is as follows: Figure 4-Figure 9 As shown in the figure, the experimental results show that UCB outperforms TS and DEG in most cases. Although TS and UCB have similar performance, UCB has better adaptability and better performance because it does not need to make strong assumptions about the distribution of the benefits of adversarial actions.
[0063] Aiming at the problems of interference from multi-interference type jammers, complex battlefield electromagnetic environment and lack of prior knowledge, the present invention constructs a complete anti-interference action library and models the benefit distribution of each anti-interference action as a device, thereby converting the confrontation problem based on the anti-interference action library into a MAB problem. Based on the working characteristics of the jammer, an online learning anti-interference framework CAMI is proposed, and an anti-interference benefit evaluation method suitable for single radar, single multi-interference type jammer and single target scenario is designed. The algorithm adopts the upper confidence bound (UCB) algorithm to solve the trade-off problem between exploration and utilization.
[0064] In the simulation experiment, the present invention implements a jammer model including six types of interference, including blocking noise amplitude modulation interference, blocking noise phase modulation interference, aiming noise phase modulation interference, intermittent sampling direct forwarding interference (ISDJ), intermittent sampling forwarding interference (ISRJ) and distance deception interference, verifying the effectiveness of CAMI-UCB in continuous multi-type interference scenarios. The experimental results show that even when a single anti-interference method fails, CAMI-UCB can still effectively combat interference and quickly solve the anti-interference action that best suits the current interference state. In comparative tests with other MAB-based algorithms, UCB shows better performance and adaptability than Attenuated Epsilon Greedy (DEG) and Thompson Sampling (TS). In summary, CAMI-UCB has the following advantages:
[0065] Combined active and passive anti-interference technology: By combining active and passive anti-interference technologies, the synergy of radar transmission technology and signal processing technology is fully utilized to obtain higher anti-interference benefits.
[0066] Adaptability to multiple interference types: It can effectively deal with multiple interference types and optimize the optimal anti-interference action at a lower computational cost and faster speed.
[0067] Embodiment 2 The present invention also provides an online anti-interference system of an intelligent radar based on MAB, including: a decision maker, an evaluator and an updater connected in sequence; the decision maker selects and outputs an anti-interference action, the anti-interference action is used to adjust the radar's transmission strategy, and the anti-interference action is stored in an anti-interference action library; the radar transmits a signal according to the adjusted transmission strategy for target detection, and after receiving the echo signal, the selected anti-interference action performs interference suppression processing on the signal; the processed signal result is transmitted to the evaluator, and the evaluator gives an anti-interference benefit score; the updater updates and optimizes the state of the decision maker according to the anti-interference benefit score, and then enters the next confrontation process. If the updater detects that the jammer changes the interference state, the current interference state E and the state parameters of the anti-interference decision algorithm are stored in the interference state memory in the form of a Key-Value pair. In a new round of confrontation, if the interference state E that appears matches the existing Key in the interference state memory, the system is initialized using its corresponding Valued, and a new round of confrontation is started; if it does not match, it is reinitialized according to the current interference state E.
[0068] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by technicians in this technical field within the essential scope of the present invention should also fall within the protection scope of the present invention.
Claims
1. An online anti-interference method for intelligent radar based on MAB, characterized in that: The specific steps include: S1, the radar and the jammer begin to confront each other, and the state parameters of the counter-interference decision algorithm are initialized according to the current interference state E; S2, build a UCB-based decision maker, which selects and outputs anti-interference actions, which are used to adjust the radar's transmission strategy; S3, the radar transmits signals according to the adjusted transmission strategy for target detection, and after receiving the echo signal, the selected anti-interference action performs interference suppression processing on the signal; S4, the processed signal result is transmitted to the anti-interference effect evaluator to evaluate the interference suppression effect and obtain the anti-interference benefit score; S5, the UCB-based updater updates and optimizes the state of the decision maker according to the anti-interference benefit score, and then enters the next confrontation process; S6, if the updater detects that the jammer changes the jamming state, the current jamming state and the corresponding state parameters of the anti-jamming decision algorithm are stored in the jamming state memory bank, otherwise the number of confrontation iterations is updated and the next confrontation is entered.
2. According to claim 1, a MAB-based intelligent radar online anti-interference method is characterized in that: The state parameters of the anti-interference decision algorithm in step S1 include: the mean anti-interference action benefit, the UCB value, the number of action selections and the number of adversarial iterations.
3. The MAB-based intelligent radar online anti-interference method according to claim 1 is characterized in that: Step S2 specifically includes the following steps: S2.1, scale the mean value of each anti-interference action benefit to interval; S2.2, calculate the UCB value, and select the action corresponding to the one with the largest UCB value as the anti-interference action output by the decision maker: ; ; ; in, is the mean benefit of the anti-interference action, It's action The number of selections, It's action No. The profit of the second choice, is the number of adversarial iterations, It's action The UCB value, is an adjustable parameter to measure uncertainty, Is a non-zero value, used to prevent the denominator from being 0. It is an anti-interference action. is the maximum value function.
4. The MAB-based intelligent radar online anti-interference method according to claim 1 is characterized in that: Step S4 specifically includes the following steps: S4.1, the number of target detections obtained after pulse compression and CFAR detection of the echo signal after interference suppression As an evaluation factor for interference suppression effect : ; S4.2, judge the authenticity of the identified target and evaluate the factors The calculation formula is as follows: ; ; in, To measure the impulse response width of the target, is the target theoretical impulse response width, is the speed of light, is the instantaneous working bandwidth of the radar, is the main lobe width error; S4.3, use the peak sidelobe ratio PSR to analyze the pulse compression results after interference suppression: ; in, is the maximum peak power in the pulse compression result, is the power of the next largest peak; S4.4, integration of three evaluation factors , and , and get the anti-interference benefit score : ; in, It is the echo signal after anti-interference action processing.
5. The MAB-based intelligent radar online anti-interference method according to claim 1 is characterized in that: Step S5 is specifically as follows: Obtaining the anti-interference benefit evaluation score Then, update the action in a step-by-step manner The mean anti-interference action benefit and the number of action selections: ; ; in, for The mean value of the anti-interference action benefit at the moment, for The anti-interference benefit score at the moment, for Moment of action The number of selections.
6. The MAB-based intelligent radar online anti-interference method according to claim 1 is characterized in that: Step S6 is as follows: if the updater detects that the jammer has changed its jamming state, the current jamming state is updated. and the corresponding anti-interference decision algorithm state parameters are stored in the interference state memory bank in the form of parameter pairs; in a new round of confrontation, if the interference state that appears matches the existing interference state in the interference state memory bank, the anti-interference decision algorithm state parameters corresponding to the interference state are used as initialization parameters, and a new round of confrontation is started; if they do not match, they are reinitialized according to the current anti-interference state.
Citation Information
Patent Citations
Through-the-wall radar imaging method based on multi-resolution fusion convolutional neural network
CN114066792A
Transmitting frequency sequence design method and device of agile coherent radar
CN118068266A
Shield pulse design method based on MAB algorithm
CN118886340A
Maneuvering target tracking method of fusion network based on MHS and GRU
CN119026089A
Method and apparatus for optimizing average bit error probability via deep multi-armed bandit in OFDM and index modulation system for low power communication
US20210160007A1