Radar adaptive anti-jamming decision method based on deep learning
By employing a radar adaptive anti-jamming decision-making method based on deep learning and multi-agent reinforcement learning, anti-jamming measures of radar systems are identified and optimized. This solves the problem that anti-jamming strategies in existing technologies cannot be adaptively optimized, and achieves fast and stable radar anti-jamming effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-19
- Publication Date
- 2026-03-27
AI Technical Summary
When faced with interference in complex electromagnetic environments, existing radar systems cannot adaptively optimize their anti-jamming strategies, and the convergence speed and stability of intelligent anti-jamming methods are limited by the accuracy of the input state.
A radar adaptive anti-jamming decision-making method based on deep learning is adopted. By acquiring the characteristic parameters of radar echo signals, deep learning algorithms are used to identify the type and characteristics of interference. Combined with multi-agent reinforcement learning algorithms, anti-jamming measures are optimized to achieve radar adaptive anti-jamming decision-making.
This improves the strategy convergence speed and stability of the radar system in complex electromagnetic environments, reduces the dependence on preset parameters and input state accuracy, and enhances the radar's anti-jamming capability and overall performance.
Smart Images

Figure CN121541150B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of radar countermeasures, in particular to a radar adaptive countermeasure decision method based on deep learning. BACKGROUND
[0002] With the complication of modern electromagnetic environment, the types, forms and intensities of various interference signals are increasingly diverse, increasing the difficulty of radar detection and tracking.
[0003] In the prior art, the anti-interference strategies mainly include active anti-interference methods and intelligent anti-interference methods.
[0004] However, the active anti-interference method mainly depends on artificial preset parameters, cannot adaptively optimize according to real-time interference changes, and the rule-driven system is unstable under low signal-to-noise ratio or unknown interference conditions; while the intelligent anti-interference method can learn the characteristics of the environment through interaction and autonomously generate optimal strategies, but the strategy convergence speed and stability are still limited by the input state accuracy. SUMMARY
[0005] Therefore, it is necessary to provide a radar adaptive countermeasure decision method based on deep learning to realize radar adaptive countermeasure decision, improve the strategy convergence speed and stability, and reduce the dependence on input state accuracy.
[0006] The radar adaptive countermeasure decision method based on deep learning comprises:
[0007] A plurality of radar echo signals are obtained, and after extracting feature parameters, class label labeling is performed, and a deep learning algorithm is used to obtain an interference class vector and an interference feature vector;
[0008] Based on the interference feature vector, the feature similarity of the radar transmitting signal and the radar echo signal in different domains is calculated respectively to obtain the risk degree in different domains; the domain with the maximum risk degree is selected as the scope of the anti-interference strategy;
[0009] Different domains are regarded as different agents; different domains are combined to obtain a plurality of transformation domains; the transformation domains are encoded to obtain a transformation domain combination vector; the transformation domain combination vector and the interference class vector are spliced to obtain an observation state vector of the agent;
[0010] The observation state vector of the agent is taken as the input of the multi-agent reinforcement learning algorithm, the agent selects the optimal measure combination of the scope in the electronic protection anti-interference measure set as the final action, and the multi-agent reinforcement learning algorithm is outputted to realize radar adaptive countermeasure decision.
[0011] In one embodiment, a plurality of radar echo signals are acquired, and after extracting feature parameters, class label annotation is performed, and a deep learning algorithm is used to obtain an interference class vector and an interference feature vector, including:
[0012] A plurality of radar echo signals are acquired, feature parameters of each radar echo signal are extracted, the extracted feature parameters are normalized, and the normalized signals are labeled with class labels, and all signals labeled with class labels are combined to form a data set;
[0013] The data set is input into a deep learning network to output an interference class vector and an interference feature vector.
[0014] In one embodiment, the feature parameters include: carrier frequency, pulse width, pulse repetition interval, pulse amplitude, and instantaneous bandwidth;
[0015] The carrier frequency is used to represent the center frequency of the pulse signal; the pulse width reflects the duration of a single radar pulse; the pulse repetition interval describes the time interval between adjacent pulses; the pulse amplitude is used to distinguish constant amplitude interference from amplitude modulation type interference; and the instantaneous bandwidth represents the bandwidth characteristics of the pulse signal in the frequency domain.
[0016] In one embodiment, the interference classes include: noise interference, false target deception jamming, decoy deception jamming, and composite jamming;
[0017] The noise interference includes sweep jamming, barrage jamming, and aiming jamming; the false target deception jamming includes dense false target jamming and range-amplitude deception jamming; the decoy deception jamming includes range decoy jamming and speed decoy jamming; and the composite jamming includes a combination of noise interference, false target deception jamming, and decoy deception jamming.
[0018] In one embodiment, the agent selects an optimal measure combination in the scope of the electronic protection anti-jamming measure set as the final action, including:
[0019] When the interference class is noise interference, the agent selects a measure combination in the scope of the electronic protection anti-jamming measure set, takes the difference in SINR before and after the measure combination is taken as a reward function, and performs cyclic iteration until the difference in SINR before and after the measure combination is taken meets a preset first condition, and the current measure combination is the optimal measure combination for noise interference, and the optimal measure combination for noise interference is taken as the final action.
[0020] When the interference category is false target deception jamming, the agent selects a scope measure combination in the electronic protection anti-jamming measure set, takes the real target recognition rate as a reward function, and performs cyclic iteration until the real target recognition rate meets a preset second condition, and the current measure combination is the optimal measure combination of the false target deception jamming, and the optimal measure combination of the false target deception jamming is taken as the final action.
[0021] When the interference category is false target deception jamming, the agent selects a scope measure combination in the electronic protection anti-jamming measure set, takes the real target recognition rate as a reward function, and performs cyclic iteration until the real target recognition rate meets a preset second condition, and the current measure combination is the optimal measure combination of the false target deception jamming, and the optimal measure combination of the false target deception jamming is taken as the final action.
[0022] When the interference category is false target deception jamming, the agent selects a scope measure combination in the electronic protection anti-jamming measure set, takes the real target recognition rate as a reward function, and performs cyclic iteration until the real target recognition rate meets a preset second condition, and the current measure combination is the optimal measure combination of the false target deception jamming, and the optimal measure combination of the false target deception jamming is taken as the final action.
[0023] In one embodiment, the different domains include: a frequency domain, a waveform domain, a space domain, and a signal processing domain.
[0024] In one embodiment, based on the interference feature vector, feature similarities of the radar transmitting signal and the radar echo signal in different domains are respectively calculated to obtain risk degrees in different domains, including:
[0025] Based on the interference feature vector, a feature similarity of a carrier frequency of the radar transmitting signal and the radar echo signal in the frequency domain is calculated, and the feature similarity of the carrier frequency is taken as the risk degree in the frequency domain.
[0026] Based on the interference feature vector, a feature similarity of a waveform correlation coefficient of the radar transmitting signal and the radar echo signal in the waveform domain is calculated, and the feature similarity of the waveform correlation coefficient is taken as the risk degree in the waveform domain.
[0027] Based on the interference feature vector, a feature similarity of an antenna beam pointing of the radar transmitting signal and the radar echo signal in the space domain is calculated, and the feature similarity of the antenna beam pointing is taken as the risk degree in the space domain.
[0028] Based on the interference feature vector, a feature similarity of a pulse amplitude of the radar transmitting signal and the radar echo signal in the signal processing domain is calculated, and the feature similarity of the pulse amplitude is taken as the risk degree in the signal processing domain.
[0029] In one embodiment, different domains are combined to obtain a plurality of transform domains; the transform domains are encoded to obtain a transform domain combination vector, including:
[0030] The frequency domain, waveform domain, space domain, and signal processing domain are combined to obtain 15 transform domains;
[0031] The 15 transform domains are binary encoded according to 1 to 15 to obtain an 8-dimensional transform domain combination vector.
[0032] In one embodiment, the electronic protection anti-jamming measure set includes 9 electronic protection anti-jamming measures, which are: inter-pulse frequency agility, sub-pulse frequency agility, LFM signal based on frequency modulation, phase encoding based signal, frequency agility signal, adaptive beamforming, sidelobe cancellation, constant false alarm detection, and pulse accumulation.
[0033] In one embodiment, when selecting the optimal measure combination of the scope, a priori knowledge base is set as a constraint for selection;
[0034] Setting the priori knowledge base includes: combining the 9 electronic protection anti-jamming measures to obtain a plurality of initial combinations; screening all initial combinations that cannot be used in cascade as the priori knowledge base.
[0035] The above radar adaptive anti-jamming decision method based on deep learning takes different domains as different agents, concatenates the transform domain combination vector and the interference category vector as the observation state vector of the agent, inputs the multi-agent reinforcement learning algorithm, and combines the deep learning model, which can not only accurately identify the interference state, but also quickly select the optimal anti-jamming transform domain through multi-domain cooperation optimization among multiple agents, and generate a multi-technology combination anti-jamming strategy. The optimization of the strategy can quickly adjust in a dynamic complex electromagnetic environment, reduce the dependence on preset parameters and input state accuracy, thereby improving the strategy convergence speed and stability, improving the reaction speed and adaptability of the system, greatly improving the anti-jamming ability and overall performance of the radar, and realizing the adaptive anti-jamming of the radar system. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 The flowchart of the radar adaptive anti-jamming decision method based on deep learning in one embodiment;
[0037] Figure 2 The architecture diagram of the Transformer-LSTM network in one embodiment;
[0038] Figure 3 The architecture diagram of the MADDPG multi-agent reinforcement learning algorithm in one embodiment;
[0039] Figure 4Fig. 1 is a schematic diagram of an architecture of a CTDE mechanism in one embodiment;
[0040] Figure 5 Fig. 1 is a schematic diagram of an architecture of a CTDE mechanism in one embodiment;
[0041] Figure 6 Fig. 1 is a schematic diagram of an architecture of a CTDE mechanism in one embodiment;
[0042] Figure 7 Fig. 1 is a schematic diagram of an architecture of a CTDE mechanism in one embodiment;
[0043] Reference Signs:
[0044] 601 first module, 602 second module, 603 third module, 604 fourth module. DETAILED DESCRIPTION
[0045] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application.
[0046] In addition, the description such as "first", "second" and the like in the present application is only for the purpose of description and should not be understood as indicating or implying the relative importance of the technical features indicated or implicitly indicating the number of technical features. Therefore, the features limited by "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "multiple groups" is at least two groups, such as two groups, three groups, etc., unless otherwise specifically limited.
[0047] In the present application, unless otherwise specifically defined and limited, the terms "connection", "fixing" and the like should be understood broadly, for example, "fixing" can be fixed connection, or detachable connection, or integral; can be mechanical connection, or electrical connection, or physical connection, or wireless communication connection; can be directly connected, or indirectly connected through an intermediate medium, can be the internal connection of two elements or the interaction relationship between two elements, unless otherwise specifically limited. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0048] In addition, the technical solutions among the various embodiments of the present application can be combined with each other, but it must be based on the implementation by the ordinary skilled in the art, and when the combination of the technical solutions appears contradictory or unimplementable, it should be considered that the combination of the technical solutions does not exist and is not within the protection scope required by the present application.
[0049] The present application provides a radar adaptive anti-jamming decision method based on deep learning, as shown in the flowchart, in one embodiment, comprising: Figure 1
[0050] Step 101, obtaining a plurality of radar echo signals, after extracting the characteristic parameters, classifying label annotation, and using deep learning algorithm to obtain the interference category vector and the interference characteristic vector.
[0051] Specifically:
[0052] Obtaining a plurality of radar echo signals, extracting the characteristic parameters of each radar echo signal, normalizing the characteristic parameters extracted for each radar echo signal, and labeling the signals after normalization (i.e., the category of active jamming) with class labels, and all signals labeled with class labels are combined into a data set (the data set can be divided into a training set and a test set in a ratio of 7:3);
[0053] The data set (the data set is a plurality of 5-dimensional vectors) is input into a deep learning network (for example, a Transformer-LSTM network) to output an interference category vector and an interference characteristic vector as an interference recognition result.
[0054] Among them:
[0055] The plurality of radar echo signals can be simulated by using MATLAB simulation software, for example, using MATLAB simulation software to generate time-domain and frequency-domain waveforms of interference signals under different characteristic parameters as a plurality of radar echo signals.
[0056] The characteristic parameters include carrier frequency, pulse width, pulse repetition interval, pulse amplitude and instantaneous bandwidth; the carrier frequency is used to represent the center frequency of the pulse signal, which is convenient for distinguishing frequency agility and frequency suppression type jamming; the pulse width reflects the duration of a single radar pulse, which has important reference value for detecting pulse envelope distortion and wide pulse jamming; the pulse repetition interval describes the time interval between adjacent pulses, which can assist in judging pulse jitter interference and pulse replay interference; the pulse amplitude is used to distinguish constant amplitude interference and amplitude modulation type interference; the instantaneous bandwidth represents the bandwidth characteristics of the pulse signal in the frequency domain, which is particularly important for identifying linear frequency modulation interference and sweep frequency interference.
[0057] Normalization includes: linear normalization, which maps feature parameters of different dimensions to the [0,1] interval, helping to eliminate the interference of differences in feature dimensions on network training. The feature vector obtained after normalization is a continuous real number, not a discretization, but rather retains the physical information of continuous values. The normalization process can be expressed as:
[0058] ;
[0059] In the formula, Normalized value For the first a In the nth sample b Original values of the feature parameters, This is the minimum value of the feature parameter across all samples. This is the maximum value of the feature parameter across all samples.
[0060] There are four types of interference: noise interference, false target deception interference, dragging deception interference, and composite interference.
[0061] Noise interference includes three types: frequency sweeping interference, jamming interference, and targeting interference. Frequency sweeping interference covers the radar's operating frequency band by linearly scanning the frequency over time, forming a continuous suppression, which can be suppressed by pulse frequency agility. Jamming interference uses broadband noise to cover multiple frequency bands simultaneously, causing a large-scale decrease in signal-to-noise ratio, which is usually suppressed by spatial filtering (such as adaptive beamforming). Targeting interference uses narrowband noise to precisely act on the radar's center frequency, suppressing it with high power concentration, which can be suppressed by frequency agility or polarization filtering.
[0062] There are two types of decoy jamming: dense decoy jamming and range-amplitude decoy jamming. Dense decoy jamming generates a large number of false targets in the range-Doppler plane, saturating the radar data processing stage. It can be identified and suppressed by waveform domain techniques such as phase coding. Range-amplitude decoy jamming creates range ambiguity by copying and delaying the target echo, causing target tracking failure. It can be detected by time-frequency analysis methods.
[0063] There are two types of dragging deception jamming: range dragging jamming and velocity dragging jamming. Range dragging jamming gradually increases the echo delay, causing the radar range tracking loop to deviate from the real target, which can be canceled by Doppler domain filtering. Velocity dragging jamming gradually changes the Doppler frequency shift, causing the velocity tracking loop to lose lock and the target to be lost, which can be canceled by joint frequency and waveform processing.
[0064] The composite jamming includes a combination of noise jamming, false target deception jamming and decoy deception jamming; the composite jamming not only suppresses radar receiving signals, but also creates a large number of false targets, causing double damage; for such multi-modal jamming, multi-domain cooperative processing in the frequency domain, the spatial domain and the waveform domain is needed to achieve comprehensive suppression effect.
[0065] The schematic diagram of the architecture of the Transformer-LSTM network is shown in FIG. 1. Figure 2 The 5-dimensional vector is input into the LSTM, and then sequentially passes through an encoder layer, a flattening layer, a fully connected layer and an output layer, and finally outputs a jamming category vector and a jamming feature vector.
[0066] The jamming category vector is a M dimensional binary coding vector (i.e., a 0-1 coding vector, each dimension corresponds to a jamming category, and the corresponding dimension is 1 if the jamming category exists, and is 0 if the jamming category does not exist), which can be specifically expressed as , M wherein n represents the number of jamming categories, for example: t At time t, there are 8 jamming categories (sweep jamming, blocking jamming, aiming jamming, dense false target jamming, range-amplitude deception jamming, range decoy jamming, speed decoy jamming and composite jamming), and the current is the third (i.e., M n = 8, m t = 3, ) jamming category vector .
[0067] The jamming feature vector refers to a vector composed of feature parameters.
[0068] In this step, feature extraction, normalization processing and label annotation are performed to achieve high-precision jamming recognition.
[0069] It should be noted that how to extract feature parameters, how to perform label annotation and how to form a data set and the Transformer-LSTM network all belong to the prior art, and will not be described here.
[0070] In step 102, based on the jamming feature vector, the feature similarity of the radar transmitting signal and the radar echo signal in different domains is calculated respectively to obtain the risk degree in different domains; the domain with the largest risk degree is selected as the scope of the anti-jamming strategy.
[0071] Specifically:
[0072] The different domains include: the frequency domain, the waveform domain, the spatial domain and the signal processing domain.
[0073] Based on the interference feature vector, the feature similarity of the carrier frequency of the radar transmitting signal and the radar echo signal (i.e. the interference signal) in the frequency domain is calculated, and the feature similarity of the carrier frequency is taken as the risk degree in the frequency domain;
[0074] Based on the interference feature vector, the feature similarity of the waveform correlation coefficient of the radar transmitting signal and the radar echo signal (i.e. the interference signal) in the waveform domain is calculated, and the feature similarity of the waveform correlation coefficient is taken as the risk degree in the waveform domain;
[0075] Based on the interference feature vector, the feature similarity of the antenna beam pointing of the radar transmitting signal and the radar echo signal (i.e. the interference signal) in the space domain is calculated, and the feature similarity of the antenna beam pointing is taken as the risk degree in the space domain;
[0076] Based on the interference feature vector, the feature similarity of the pulse amplitude of the radar transmitting signal and the radar echo signal (i.e. the interference signal) in the signal processing domain is calculated, and the feature similarity of the pulse amplitude is taken as the risk degree in the signal processing domain;
[0077] The risk degrees in the frequency domain, the waveform domain, the space domain and the signal processing domain are compared, when the domain corresponding to the maximum risk degree is only one, the domain is taken as the scope of the anti-interference strategy, and when the domain corresponding to the maximum risk degree is more than two, the two or more domains are taken as the scope of the anti-interference strategy (this is a joint domain decision process).
[0078] Wherein:
[0079] The feature similarity is measured by the distance between the sample features, and can be calculated by using Euclidean distance divided by the feature number in the domain:
[0080] ;
[0081] In the formula, is the sample is the sample The feature similarity in the first domain, the smaller the value is, the higher the similarity degree is; is the first feature number; is the total number of features selected in a domain; is the sample is the first feature number in the first domain; is the sample is the first feature number in the first domain; is the feature number selected by the sample in the first domain.
[0082] In this step, the frequency domain selects the carrier frequency as the feature, the waveform domain selects the waveform correlation coefficient (Corr) as the feature, the spatial domain selects the antenna beam pointing as the feature, and the signal processing domain selects the pulse amplitude as the feature. Not only can the interference be applied in a certain domain (frequency domain, waveform domain, spatial domain, signal processing domain), but also the working parameters of the radar in these domains are "aligned" to improve the interference effect of the interference signal on the radar.
[0083] On this basis, the feature similarity is calculated to quantify the risk degree of the current interference in each domain, reflect the urgency of the interference measure demand in the domain, and if the feature similarity is smaller and the domain of the maximum feature similarity is more (the number of accurately matched domains is more), the risk degree of the interference is greater, the domain alignment degree is higher, the effective energy of the interference signal at the radar end is stronger, the radar performance decreases more obviously, and the demand for anti-jamming action in this domain is greater, thereby determining the most targeted defense domain and realizing domain optimization. At the same time, by determining the anti-jamming in which domain is most effective, the selection range of the subsequent selection from the electronic protection anti-jamming measure set (all ECCM technologies) can be reduced, thereby reducing the algorithm calculation amount, speeding up the decision-making speed, and reducing the risk of selecting the wrong electronic protection anti-jamming measure.
[0084] It should be noted that the frequency domain, the waveform domain, the spatial domain, and the signal processing domain all belong to the prior art and will not be described here.
[0085] Step 103, different domains are taken as different agents; different domains are combined to obtain a plurality of transformation domains; the transformation domains are encoded to obtain a transformation domain combination vector; and the transformation domain combination vector and the interference category vector are spliced to obtain an observation state vector of the agent.
[0086] Specifically:
[0087] The frequency domain, the waveform domain, the spatial domain, and the signal processing domain are taken as different agents, and there are a total of 4 agents, i.e. ;
[0088] The frequency domain, the waveform domain, the spatial domain, and the signal processing domain are combined to obtain 15 transformation domains, i.e. ;
[0089] The 15 transformation domains are binary coded according to 1 to 15 to obtain a transformation domain combination vector.
[0090] The transformation domain combination vector and the interference category vector are spliced to obtain an observation state vector of the agent.
[0091] Among them:
[0092] The transform domain combination vector is a 4-dimensional binary encoding vector (i.e., a 0-1 encoding vector), for example: t At the moment, there are 15 transform domains, and the current one is the third one (i.e., T = 3), the transform domain combination vector is v . .
[0093] The observation state vector of the agent is a 12-dimensional binary encoding vector (i.e., a 0-1 encoding vector), which is obtained by splicing the transform domain combination vector and the interference category vector, for example, t At the moment, the transform domain combination vector is , the interference category vector is , and the observation state vector of the first agent is n . .
[0094] In this step, the transform domain combination vector and the interference category vector are spliced to obtain the observation state vector of the agent, which is used as the input of the multi-agent reinforcement learning algorithm.
[0095] In step 104, the observation state vector of the agent is used as the input of the multi-agent reinforcement learning algorithm, the agent selects the optimal measure combination of the scope in the electronic protection anti-jamming measure set as the final action, and the multi-agent reinforcement learning algorithm outputs the final action to realize radar adaptive anti-jamming decision.
[0096] Specifically:
[0097] The observation state vector of the agent is used as the input of the multi-agent reinforcement learning algorithm;
[0098] The electronic protection anti-jamming measure set includes 9 electronic protection anti-jamming measures, which are: inter-pulse frequency agility, sub-pulse frequency agility, LFM signal based on frequency modulation, phase encoding based signal, frequency agility signal, adaptive beam forming, sidelobe cancellation, constant false alarm detection, and pulse accumulation;
[0099] The agent selects the measure combination of the scope in the electronic protection anti-jamming measure set, sets the prior knowledge base as the constraint of the selection, obtains the optimal measure combination, and uses the optimal measure combination as the final action;
[0100] The final action is output by the multi-agent reinforcement learning algorithm to realize radar adaptive anti-jamming decision.
[0101] More specifically:
[0102] The observation state vector of the agent is used as the input of the MADDPG multi-agent reinforcement learning algorithm;
[0103] The electronic counter-countermeasure (ECCM) set includes nine electronic counter-countermeasures, namely, pulse-to-pulse frequency agility, sub-pulse frequency agility, frequency modulation based LFM signal, phase encoding based signal, frequency agility signal, adaptive beam forming, sidelobe cancellation, constant false alarm detection, and pulse accumulation;
[0104] When the interference type is noise jamming, the agent selects a scope measure combination from the electronic counter-countermeasure set, sets a priori knowledge base (combines the nine electronic counter-countermeasures to obtain a plurality of initial combinations; screens all initial combinations that cannot be used in cascade as the priori knowledge base), as a constraint for selection (if the selected measure combination exists in the priori knowledge base, the action is judged to be illegal, the measure combination is reselected, and until the current measure combination does not exist in the priori knowledge base, the action is judged to be legal, so as to ensure the rationality of the output action), to obtain a legal combination; takes the difference between the SINR before and after the legal combination is taken as a reward function (SINR ), and performs cyclic iteration until the difference between the SINR before and after the legal combination is taken meets a preset first condition, and the current legal combination is the optimal measure combination for noise jamming, and the optimal measure combination for noise jamming is taken as the final action;
[0105] When the interference type is false target deception jamming, the agent selects a scope measure combination from the electronic counter-countermeasure set, sets a priori knowledge base (combines the nine electronic counter-countermeasures to obtain a plurality of initial combinations; screens all initial combinations that cannot be used in cascade as the priori knowledge base), as a constraint for selection (if the selected measure combination exists in the priori knowledge base, the action is judged to be illegal, the measure combination is reselected, and until the current measure combination does not exist in the priori knowledge base, the action is judged to be legal, so as to ensure the rationality of the output action), to obtain a legal combination; takes the real target recognition rate (i.e., the ratio of the number of real targets actually detected by the radar to the total number of targets detected by the system) as a reward function (SINR ), and performs cyclic iteration until the real target recognition rate meets a preset second condition, and the current measure combination is the optimal measure combination for false target deception jamming, and the optimal measure combination for false target deception jamming is taken as the final action;
[0106] When the interference category is the towed deception jamming, the intelligent agent selects the scope measure combination in the electronic protection anti-jamming measure set, sets the prior knowledge base (combines the 9 electronic protection anti-jamming measures to obtain a plurality of initial combinations; screens all initial combinations that cannot be used in cascade as the prior knowledge base), as the constraint of selection (if the selected measure combination exists in the prior knowledge base, it is judged that the action is illegal, the measure combination is reselected, and until the current measure combination does not exist in the prior knowledge base, it is judged that the action is legal, so as to ensure the rationality of the output action), to obtain a legal combination; the negative value of the tracking accuracy error is taken as a reward function ( ), and the loop iteration is performed until the negative value of the tracking accuracy error meets a preset third condition, and the current measure combination is the optimal measure combination of the towed deception jamming, and the optimal measure combination of the towed deception jamming is taken as the final action.
[0107] When the interference category is the composite jamming, the intelligent agent selects the scope measure combination in the electronic protection anti-jamming measure set, sets the prior knowledge base (combines the 9 electronic protection anti-jamming measures to obtain a plurality of initial combinations; screens all initial combinations that cannot be used in cascade as the prior knowledge base), as the constraint of selection (if the selected measure combination exists in the prior knowledge base, it is judged that the action is illegal, the measure combination is reselected, and until the current measure combination does not exist in the prior knowledge base, it is judged that the action is legal, so as to ensure the rationality of the output action), to obtain a legal combination; the weighted sum of the difference between the SINR before and after the measure combination is taken, the real target recognition rate and the negative value of the tracking accuracy error (how to take each weight value is the prior art, and specific adjustment can be made according to the interference type contained in the composite jamming, which will not be described here) is taken as a reward function ( ), and the loop iteration is performed until the weighted sum meets a preset fourth condition, and the current measure combination is the optimal measure combination of the composite jamming, and the optimal measure combination of the composite jamming is taken as the final action.
[0108] The final action is output by the MADDPG multi-agent reinforcement learning algorithm and is executed by the radar system, so as to realize the radar adaptive anti-jamming decision.
[0109] Among them:
[0110] The electronic protection anti-jamming measures contained in each domain are shown in Table 1:
[0111] Table 1: Electronic protection anti-jamming measures
[0112]
[0113] The final action (i.e. the optimal measure combination of the interference) output by the MADDPG multi-agent reinforcement learning algorithm is a WA binary encoding vector (i.e., a 0-1 encoding vector, each dimension corresponds to an electronic protection anti-jamming measure, and the corresponding dimension is 1 when the electronic protection anti-jamming measure is selected, and 0 when it is not selected), which can be specifically expressed as , W represents the number of electronic protection anti-jamming measures selected in the scope, for example: t At the moment, there are 4 electronic protection anti-jamming measures in the scope and the 2nd one is selected (i.e. W = 4, ) The final action is .
[0114] In this step, in the selected scope, the action closed-loop optimization is carried out in combination with the multi-agent reinforcement learning algorithm and the electronic protection anti-jamming measure set, so as to realize the radar anti-jamming decision.
[0115] As shown in the architecture diagram of the MADDPG multi-agent reinforcement learning algorithm Figure 3 The whole is composed of environment, multiple agents (taking i agents as an example, only agent 1, agent 2, agent 3 and agent i are shown in the figure; all agents have the same structure, and a double network design of "online network + target network" is adopted) and experience replay pool three parts constitute a closed loop interaction. The complex electromagnetic environment in which the radar is located will output the corresponding state perceived by each agent, i.e., the fusion representation of the current environment information, which is a 12-dimensional observation state vector composed of an 8-dimensional interference category vector and a 4-dimensional transform domain combined vector, to reflect "current interference type + scope to be defended", and each agent outputs action according to the state, i.e., a binary encoding combined vector is selected from 9 electronic protection anti-jamming measures, which represents the anti-jamming operation that should be performed in this domain, such as inter-pulse frequency agility, adaptive beam forming, etc. The processes of the first three agents in the figure are the same: the online policy network outputs action in real time according to the current state (where is the online policy network parameter), and the target policy network outputs action in real time according to the next state (where is the target policy network parameter), which is used for subsequent target value calculation; the corresponding online value network evaluates the online value of the state-action pair (where is the online value network parameter), and the target value network provides stable value estimation and outputs the target value. Agent i further refines the training process: after the online policy network outputs , the OU noise is superimposed to form the final action To balance exploration and utilization; online value output by online value networks The target value output by the target value network (in For the target value network parameters, Output for other intelligent agents The summation of the two values (the summation of the ... GB Update online network parameters Then, soft updates are used to slowly synchronize the online network parameters to the target policy network parameters and the target value network parameters to maintain training stability. The experience replay pool is used to store experience tuples generated by the agent's interactions with the environment. ,in, s This is the current state. s ′ represents the next state. r The reward for environmental feedback (such as the improvement in SINR after interference immunity) is input into the agent as... r i ), a Actions performed by the intelligent agent d The flag indicates whether a round has ended; during training, small batches of samples are randomly sampled for network updates to break data correlation and improve learning efficiency. All parameters in the figure collectively embody the collaborative decision-making and closed-loop optimization mechanism of multi-agent systems in various radar anti-jamming domains. Stable training is achieved through a dual-network structure and experience replay, ultimately outputting a combination of anti-jamming measures adapted to complex electromagnetic environments.
[0116] Because different anti-jamming technologies (electronic protection anti-jamming measures) correspond to different signal processing links and physical principles, their effects may reinforce or cancel each other out, or fail under specific environments. Therefore, it is difficult for a single agent to simultaneously optimize the use strategy of multiple technologies. The MADDPG multi-agent reinforcement learning algorithm enables each agent to focus on a single domain and, during training, utilizes a centralized training and distributed execution (CTDE) mechanism (such as... Figure 4 As shown, this illustrates the collaborative relationship between multiple agent networks (Actor network) and a centralized Critic network during the training phase to achieve policy coordination.
[0117] The prior knowledge base is set (9 electronic protection anti-interference measures are combined to obtain a plurality of initial combinations; all initial combinations that cannot be used in cascade are screened out as the prior knowledge base), used as a constraint for selection (if the selected measure combination exists in the prior knowledge base, it is judged that the action is illegal, the measure combination is reselected, and until the current measure combination does not exist in the prior knowledge base, it is judged that the action is legal, so as to ensure the rationality of the output action), a legal combination is obtained, which can avoid the problem that part of the electronic protection anti-interference measures in the domain cannot be used at the same time in the anti-interference process, and avoid the failure of anti-interference effect or radar system failure caused by ECCM compatibility conflict; for example, the pulse-to-pulse agile waveform in the frequency domain and the linear frequency modulation (LFM) signal electronic protection anti-interference measure in the waveform domain cannot be used at the same time.
[0118] For noise jamming, the core mechanism is to reduce the ratio of "useful signal" to "interference + background noise" received by the detection system by superimposing random noise, and ultimately to cause the useful signal to be submerged; the signal-to-interference-and-noise ratio (SINR) can directly quantify the damage degree of noise jamming to the "signal quality", and the difference between the SINRs before and after the legal combination is taken as the reward function, so as to most directly reflect whether the anti-interference strategy is effective in resisting the "signal submersion" problem and judging the effectiveness of the radar anti-interference strategy.
[0119] For false target deception jamming, the core mechanism is to generate a large number of false target signals by releasing radar decoys and faking false echoes, which will interfere with the target recognition process of the detection system and make it difficult to distinguish "real target" from "false target", and ultimately cause the system to have problems such as false target tracking and real target missing; the real target recognition rate (i.e. the ratio of the number of real targets actually detected by the radar to the total number of targets detected by the system) is used as the reward function, which can directly reflect the strength of the detection system in distinguishing real targets in a mixed scene of real and false targets, quantify the anti-interference effect of this type of jamming, and further accurately judge the effectiveness of the anti-interference strategy.
[0120] For the drag deception jamming, the core mechanism is to release gradual false signals to lure the tracking module of the detection system to deviate, causing the tracking of the real target to deviate; the negative value of the tracking accuracy error is used as the reward function, which can quantify the deviation between the tracking position of the system and the actual position of the real target, and the smaller the error, the more effective the anti-interference; this index directly corresponds to the essence of "tracking deviation" of the jamming, can reflect the ability of the system to maintain the tracking accuracy of the real target, and can be converted into a reward value consistent with the logic of "the higher the reward, the better the anti-interference" after being designed as a negative reward.
[0121] For composite jamming, its core mechanism is to cause damage to the system from multiple dimensions such as signal, identification, tracking, etc. The weighted sum of the difference in SINR before and after taking measures, the real target recognition rate, and the negative value of the tracking accuracy error is taken as the reward function, which can reflect the priority difference of different tasks in each dimension through weight setting, comprehensively quantify the overall anti-jamming capability of the system under the simultaneous action of multiple jammers, avoid the one-sidedness of a single index, and comprehensively evaluate the anti-jamming effect of this multi-dimensional damage.
[0122] It should be noted that the MADDPG multi-agent reinforcement learning algorithm, the electronic protection anti-jamming measure, how to select the initial combination that cannot be used in cascade, how to calculate the SINR difference, how to get the number of real targets detected by the radar, how to get the total number of targets detected by the system, how to get the tracking accuracy error, the first condition, the second condition, the third condition and the fourth condition are all prior art, and will not be repeated here.
[0123] In the present application, as shown in Figure 5 First, interference identification is performed, and 5-dimensional features are extracted, outputting an interference category vector and an interference feature vector. Then, domain optimization is performed, the similarity of each domain is calculated, and the optimal scope is determined. Subsequently, combination encoding is performed, each domain is taken as an agent, and the observation state vector of the agent is obtained. Finally, policy output is performed, a multi-agent reinforcement learning framework is adopted, anti-jamming technology combinations are optimized within the scope, prior knowledge base is introduced to ensure policy compatibility, and a reward function is constructed through multiple indexes to realize policy closed-loop optimization.
[0124] The above-mentioned radar adaptive anti-jamming decision method based on deep learning takes different domains as different agents, splices the transformation domain combination vector and the interference category vector as the observation state vector of the agent, inputs the multi-agent reinforcement learning algorithm, and combines the deep learning model, which can not only accurately identify the interference state, but also quickly select the optimal anti-jamming transformation domain through multi-domain cooperation optimization among multiple agents, and generate a multi-technology combination anti-jamming strategy. The optimization of the strategy can quickly adjust in a dynamic complex electromagnetic environment, reduce the dependence on preset parameters and input state accuracy, thereby improving the strategy convergence speed and stability, improving the reaction speed and adaptability of the system, greatly improving the anti-jamming capability and overall performance of the radar, and realizing the adaptive anti-jamming of the radar system.
[0125] Specifically, the present application has the following beneficial effects:
[0126] 1) Unlike the traditional method of relying only on simple energy detection or shallow classifier, the application maps the scope of radar anti-jamming core to independent agents, each agent focuses on the optimization of anti-jamming measures in its own domain, while converting the actual feature content of the transform domain combination and interference category into vector form, and constructing the observation state space with the splicing result of the vector as the input of the multi-agent reinforcement learning algorithm to realize the fusion representation of multi-domain state information, instead of using a single-dimensional parameter or flattened features as input, thereby fundamentally overcoming the limitations of the preset parameter strategy of the traditional method; On this basis, combined with a deep learning model (which is a prior art), complex spatio-temporal features can be accurately captured, automatically extracted and effectively processed from radar echo signals, not only can effectively process high-dimensional feature signals of long time series, taking into account the extraction of local time dependence and global pattern, but also can be targeted to cope with the scene where multiple interference signals act simultaneously, even in a complex electromagnetic environment with low signal-to-noise ratio and multiple interference types overlapping, it can still maintain a high interference recognition accuracy, realize high-precision interference recognition and pattern classification, significantly improve the robustness, reliability and adaptive anti-jamming capability of the system, making the radar perform more stably in complex electromagnetic environments, and quickly adapt to environmental changes and effectively cope with changing interference, ensuring the efficient operation of the radar system.
[0127] 2) Compared with the traditional rule-driven method, the features of the application work together to overcome the defects of the prior art that the agent reinforcement learning can only optimize a single-dimensional anti-jamming strategy, allowing different domain agents to work collaboratively through centralized training and decentralized execution mode, not only solving the compatibility conflict problem between multi-domain measures, but also reducing the decision space through the strategy of "domain risk degree priority scope selection", greatly improving the algorithm convergence speed and decision accuracy; At the same time, the application can provide more rich feature expression and related information, so as to more accurately capture the multi-dimensional dynamic characteristics of the interference signal, improve the interference strategy convergence speed and flexibility, and directly further promote the significant improvement of the system robustness, reliability and adaptive anti-jamming capability, ensure the interference effect, reduce manual intervention, improve the survivability and stability of detection and tracking, and enable the radar system to quickly adapt to environmental changes and effectively cope with changing interference, ensuring stable performance in complex electromagnetic environments and efficient operation.
[0128] The application is suitable for radar countermeasure autonomous decision-making technical field, especially radar countermeasure scene in complex electromagnetic environment (such as radar anti-jamming system), and has wide application prospect.
[0129] It should be understood that, although Figure 1The steps in the flowchart of FIG. 1 are displayed in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise explicitly stated herein, there is no strict order limitation for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 At least part of the steps in FIG. 1 can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be alternately executed with other steps or at least part of the sub-steps or stages of other steps.
[0130] The application also provides a radar adaptive anti-jamming decision device based on deep learning, as shown in FIG. 2, which comprises a first module 201, a second module 202, a third module 203, and a fourth module 204, wherein: Figure 6
[0131] The first module 201 is configured to obtain a plurality of radar echo signals, label the category after extracting the feature parameters, and obtain the interference category vector and the interference feature vector by using a deep learning algorithm (i.e., an interference identification module).
[0132] The second module 202 is configured to calculate the feature similarity of the radar transmitting signal and the radar echo signal in different domains based on the interference feature vector, obtain the risk degree in different domains, and select the domain with the largest risk degree as the scope of the anti-jamming strategy (i.e., a domain optimization module).
[0133] The third module 203 is configured to regard different domains as different agents, combine different domains to obtain a plurality of transformation domains, encode the transformation domains to obtain a transformation domain combination vector, and splice the transformation domain combination vector and the interference category vector to obtain an observation state vector of the agent (i.e., an encoding module).
[0134] The fourth module 204 is configured to take the observation state vector of the agent as the input of a multi-agent reinforcement learning algorithm, select the optimal measure combination of the scope in the electronic protection anti-jamming measure set as the action output by the multi-agent reinforcement learning algorithm, so as to realize radar adaptive anti-jamming decision (i.e., an ECCM selection module).
[0135] The specific limitations of the radar adaptive anti-jamming decision device based on deep learning can refer to the limitations of the radar adaptive anti-jamming decision method based on deep learning in the above, which will not be repeated here. Each module in the above device can be realized by software, hardware and their combination in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in the form of software, so as to call and execute the operations corresponding to the above modules by the processor.
[0136] In one embodiment, a computer device, which can be a terminal, is provided, and an internal structure diagram of the computer device can be as shown in Figure 7 The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is used to communicate with external terminals through network connection. The computer program is executed by the processor to implement a radar adaptive anti-jamming decision method based on deep learning. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.
[0137] Those skilled in the art can understand that Figure 7 The structure shown in the above
[0138] In one embodiment, a computer device is provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the method in the above embodiments.
[0139] In one embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps of the method in the above embodiments.
[0140] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, databases, or other media in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct Rambus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0141] The contents not described in detail in the specification belong to the prior art known to those skilled in the art.
[0142] The technical features of the above embodiments can be combined in any way. In order to make the description simple, not all possible combinations of the technical features in the above embodiments are described, but as long as the combination of the technical features does not exist, it should be considered as the scope of the specification.
[0143] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the present application. It should be pointed out that for those skilled in the art, without departing from the concept of the present application, some modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended application documents.
Claims
1. A radar adaptive anti-jamming decision-making method based on deep learning, characterized in that, include: Multiple radar echo signals were acquired, and after extracting feature parameters, category labels were applied. Then, a deep learning algorithm was used to obtain the interference category vector and the interference feature vector. Based on the interference feature vector, the feature similarity between the radar transmitted signal and the radar echo signal in different domains is calculated to obtain the risk level in different domains; the domain with the highest risk level is selected as the domain of application of the anti-jamming strategy. Different domains are treated as different agents; different domains are combined to obtain multiple transform domains; the transform domains are encoded to obtain a transform domain combination vector; the transform domain combination vector and the interference category vector are concatenated to obtain the agent's observation state vector; Using the observation state vector of the agent as the input of the multi-agent reinforcement learning algorithm, the agent selects the optimal combination of measures in the scope of action from the set of electronic protection and anti-jamming measures as the final action, and outputs it by the multi-agent reinforcement learning algorithm to realize radar adaptive anti-jamming decision-making. The types of interference include: noise interference, decoy target deception interference, dragging deception interference, and composite interference; Noise interference includes: frequency sweeping interference, jamming interference, and aiming interference; false target deception interference includes: dense false target interference and range-amplitude deception interference; drag deception interference includes: range drag interference and velocity drag interference; composite interference includes: a combination of noise interference, false target deception interference, and drag deception interference. The agent selects the optimal combination of measures within its scope from the set of electronic protection and anti-interference measures as its final action, including: When the interference category is noise interference, the agent selects a combination of measures within the scope of the electronic protection anti-interference measures set, uses the SINR difference before and after taking the measure combination as the reward function, and performs iterative loops until the SINR difference before and after taking the measure combination meets the preset first condition. The current measure combination is the optimal measure combination for noise interference, and the optimal measure combination for noise interference is taken as the final action. When the interference category is false target deception interference, the agent selects the combination of measures within the scope of the electronic protection anti-interference measures set, uses the real target recognition rate as the reward function, and performs iterative loops until the real target recognition rate meets the preset second condition. The current combination of measures is the optimal combination of measures for false target deception interference, and the optimal combination of measures for false target deception interference is used as the final action. When the interference category is dragging deception interference, the agent selects a combination of measures within the scope of the electronic protection anti-interference measures set, uses the negative value of the tracking accuracy error as the reward function, and performs iterative loops until the negative value of the tracking accuracy error meets the preset third condition. The current combination of measures is the optimal combination of measures for dragging deception interference, and the optimal combination of measures for dragging deception interference is used as the final action. When the interference category is complex interference, the agent selects a combination of measures within the scope of the electronic protection anti-interference measures set. The reward function is the weighted sum of the SINR difference before and after taking the measure combination, the true target recognition rate, and the negative value of the tracking accuracy error. The process is iterated until the weighted sum meets the preset fourth condition. The current measure combination is the optimal measure combination for complex interference, and the optimal measure combination for complex interference is taken as the final action.
2. The radar adaptive anti-jamming decision-making method based on deep learning according to claim 1, characterized in that, Multiple radar echo signals were acquired, and after extracting feature parameters, category labels were applied. A deep learning algorithm was then used to obtain the interference category vector and interference feature vector, including: Multiple radar echo signals are acquired, feature parameters of each radar echo signal are extracted, the extracted feature parameters are normalized, and the normalized feature parameters are labeled with category labels. All feature parameters labeled with category labels are combined into a dataset. The dataset is input into a deep learning network to output a disturbance category vector and a disturbance feature vector.
3. The radar adaptive anti-jamming decision-making method based on deep learning according to claim 2, characterized in that, Characteristic parameters include: carrier frequency, pulse width, pulse repetition interval, pulse amplitude, and instantaneous bandwidth; The carrier frequency is used to characterize the center frequency of the pulse signal; the pulse width reflects the duration of a single radar pulse; the pulse repetition interval describes the time interval between adjacent pulses; the pulse amplitude is used to distinguish between constant amplitude interference and amplitude modulation interference; and the instantaneous bandwidth characterizes the bandwidth characteristics of the pulse signal in the frequency domain.
4. The radar adaptive anti-jamming decision-making method based on deep learning according to any one of claims 1 to 3, characterized in that, Different domains include: frequency domain, waveform domain, spatial domain, and signal processing domain.
5. The radar adaptive anti-jamming decision-making method based on deep learning according to claim 4, characterized in that, Based on the interference feature vector, the feature similarity between the radar transmitted signal and the radar echo signal in different domains is calculated to obtain the risk level in different domains, including: Based on the interference feature vector, the feature similarity between the radar transmitted signal and the radar echo signal in the frequency domain is calculated, and the feature similarity of the carrier frequency is used as the risk level in the frequency domain. Based on the interference feature vector, the feature similarity of the waveform correlation coefficient between the radar transmitted signal and the radar echo signal in the waveform domain is calculated, and the feature similarity of the waveform correlation coefficient is used as the risk level in the waveform domain. Based on the interference feature vector, the feature similarity of the antenna beam pointing of the radar transmitted signal and the radar echo signal in the airspace is calculated, and the feature similarity of the antenna beam pointing is used as the risk level in the airspace. Based on the interference feature vector, the characteristic similarity of the pulse amplitude of the radar transmitted signal and the radar echo signal in the signal processing domain is calculated, and the characteristic similarity of the pulse amplitude is used as the risk level in the signal processing domain.
6. The radar adaptive anti-jamming decision-making method based on deep learning according to any one of claims 1 to 3, characterized in that, By combining different domains, multiple transformation domains can be obtained; Encoding the transform domain yields a combined transform domain vector, including: By combining the frequency domain, waveform domain, spatial domain, and signal processing domain, 15 transform domains are obtained; The 15 transform domains are encoded in binary according to the numbers 1 to 15 to obtain an 8-dimensional transform domain combination vector.
7. The radar adaptive anti-jamming decision-making method based on deep learning according to any one of claims 1 to 3, characterized in that, The set of electronic protection and anti-interference measures includes nine types of electronic protection and anti-interference measures, namely: inter-pulse frequency agility, sub-pulse frequency agility, frequency modulation-based LFM signal, phase-coded signal, frequency agility signal, adaptive beamforming, sidelobe cancellation, constant false alarm rate detection, and pulse accumulation.
8. The radar adaptive anti-jamming decision-making method based on deep learning according to claim 7, characterized in that, When selecting the optimal combination of measures for a given scope, a prior knowledge base is set up as a constraint for the selection. Setting up the prior knowledge base includes: combining nine electronic protection and anti-interference measures to obtain multiple initial combinations; and filtering out all initial combinations that cannot be cascaded as the prior knowledge base.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle action decision-making method and device based on reinforcement learning
CN111708355A
Radar anti-interference intelligent decision-making method based on reinforcement learning
CN113625233A