Radar adaptive anti-interference decision method based on deep learning

By proposing a radar adaptive anti-jamming decision-making method based on deep learning and multi-agent reinforcement learning, the problem of unstable anti-jamming strategies in existing radar systems in complex electromagnetic environments is solved, and rapid adaptive anti-jamming is achieved, thereby improving the radar's anti-jamming capability and overall performance.

CN121541150AActive Publication Date: 2026-02-17NAT UNIV OF DEFENSE TECH
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202610064325.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-19
Publication Date
2026-02-17
Estimated Expiration
2046-01-19

AI Technical Summary

Technical Problem

When faced with interference in complex electromagnetic environments, existing radar systems cannot adaptively optimize their anti-jamming strategies, and they are unstable under low signal-to-noise ratio or unknown interference conditions. The convergence speed and stability of existing intelligent anti-jamming methods are limited by the accuracy of the input state.

Method used

A radar adaptive anti-jamming decision-making method based on deep learning is adopted. By acquiring the characteristic parameters of radar echo signals, deep learning algorithms are used to identify the type and characteristics of interference. Combined with multi-agent reinforcement learning algorithms, anti-jamming strategies are optimized in different domains, generating multi-technology combined anti-jamming strategies, reducing the dependence on preset parameters and the accuracy of input states.

Benefits of technology

This technology enables rapid adaptive anti-jamming of radar systems in dynamic and complex electromagnetic environments, improves strategy convergence speed and stability, and enhances the radar's anti-jamming capability and overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121541150A_ABST
    Figure CN121541150A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of radar anti-interference, and relates to a radar adaptive anti-interference decision-making method based on deep learning, and the method comprises the steps: obtaining a plurality of radar echo signals, and obtaining an interference type vector and an interference feature vector; on the basis of the interference feature vectors, feature similarities of the radar transmitting signals and the radar echo signals in different domains are calculated respectively, and the action range of an anti-interference strategy is obtained; taking different domains as different agents; combining and coding different domains to obtain a transform domain combination vector, and combining the interference category vector to obtain an observation state vector of the intelligent agent; and the observation state vector of the agent is used as the input of a multi-agent reinforcement learning algorithm, the agent selects the final action of the action range, and the final action is output by the multi-agent reinforcement learning algorithm, so that a radar adaptive anti-interference decision is realized. According to the invention, the radar adaptive anti-interference decision can be realized, and the strategy convergence speed and stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of radar countermeasures, in particular to a radar adaptive countermeasure decision method based on deep learning. BACKGROUND

[0002] With the complication of modern electromagnetic environment, the types, forms and intensities of various interference signals are increasingly diverse, increasing the difficulty of radar detection and tracking.

[0003] In the prior art, the anti-interference strategies mainly include active anti-interference methods and intelligent anti-interference methods.

[0004] However, the active anti-interference method mainly depends on artificial preset parameters and cannot adaptively optimize according to real-time interference changes, and the rule-driven system is unstable under low signal-to-noise ratio or unknown interference conditions; while the intelligent anti-interference method can learn the characteristics of the environment through interaction and autonomously generate optimal strategies, but the strategy convergence speed and stability are still limited by the input state accuracy. SUMMARY

[0005] Therefore, it is necessary to provide a radar adaptive countermeasure decision method based on deep learning to realize radar adaptive countermeasure decision, improve the strategy convergence speed and stability, and reduce the dependence on input state accuracy.

[0006] The radar adaptive countermeasure decision method based on deep learning comprises: a plurality of radar echo signals are obtained, and after extracting feature parameters, class label labeling is performed, and a deep learning algorithm is used to obtain an interference category vector and an interference feature vector; based on the interference feature vector, the feature similarity of the radar transmitting signal and the radar echo signal in different domains is calculated respectively to obtain the risk degree in different domains; the domain with the maximum risk degree is selected as the scope of the anti-interference strategy; different domains are regarded as different agents; different domains are combined to obtain a plurality of transformation domains; the transformation domains are encoded to obtain a transformation domain combination vector; the transformation domain combination vector and the interference category vector are spliced to obtain an observation state vector of the agent; the observation state vector of the agent is taken as the input of the multi-agent reinforcement learning algorithm, the agent selects the optimal measure combination of the scope in the electronic protection anti-interference measure set as the final action, and the multi-agent reinforcement learning algorithm is output to realize radar adaptive countermeasure decision.

[0007] In one embodiment, obtaining a plurality of radar echo signals, performing class label labeling after extracting feature parameters, and using a deep learning algorithm to obtain an interference category vector and an interference feature vector comprises: acquire a plurality of radar echo signals, extract a characteristic parameter of each radar echo signal, normalize the extracted characteristic parameter, and label the normalized signal with a category label, and all signals labeled with a category label form a data set; input the data set into a deep learning network to output an interference category vector and an interference feature vector.

[0008] In one embodiment, the characteristic parameters include: carrier frequency, pulse width, pulse repetition interval, pulse amplitude, and instantaneous bandwidth; The carrier frequency is used to represent the center frequency of the pulse signal; the pulse width reflects the duration of a single radar pulse; the pulse repetition interval describes the time interval between adjacent pulses; the pulse amplitude is used to distinguish constant amplitude interference from amplitude modulation type interference; and the instantaneous bandwidth represents the bandwidth characteristics of the pulse signal in the frequency domain.

[0009] In one embodiment, the interference categories include: noise interference, false target deception jamming, decoy deception jamming, and composite jamming; The noise interference includes: sweep jamming, barrage jamming, and aiming jamming; the false target deception jamming includes: dense false target jamming and range-amplitude deception jamming; the decoy deception jamming includes: range decoy jamming and speed decoy jamming; and the composite jamming includes: a combination of noise interference, false target deception jamming, and decoy deception jamming.

[0010] In one embodiment, the agent selects an optimal combination of measures in the scope of the electronic protection anti-jamming measure set as the final action, including: When the interference category is noise interference, the agent selects a combination of measures in the scope of the electronic protection anti-jamming measure set, takes the difference in SINR before and after the combination of measures as the reward function, and performs cyclic iteration until the difference in SINR before and after the combination of measures meets a preset first condition, at which point the current combination of measures is the optimal combination of measures for noise interference, and the optimal combination of measures for noise interference is taken as the final action; When the interference category is false target deception jamming, the agent selects a combination of measures in the scope of the electronic protection anti-jamming measure set, takes the real target recognition rate as the reward function, and performs cyclic iteration until the real target recognition rate meets a preset second condition, at which point the current combination of measures is the optimal combination of measures for false target deception jamming, and the optimal combination of measures for false target deception jamming is taken as the final action; When the interference category is decoy deception jamming, the agent selects a combination of measures in the scope of the electronic protection anti-jamming measure set, takes the negative value of the tracking accuracy error as the reward function, and performs cyclic iteration until the negative value of the tracking accuracy error meets a preset third condition, at which point the current combination of measures is the optimal combination of measures for decoy deception jamming, and the optimal combination of measures for decoy deception jamming is taken as the final action. When the interference type is the composite interference, the agent selects a scope measure combination in the electronic protection anti-interference measure set, takes a weighted sum of a difference in SINR before and after the measure combination, a real target recognition rate and a negative value of a tracking accuracy error as a reward function, and performs a loop iteration until the weighted sum satisfies a preset fourth condition, and the current measure combination is the optimal measure combination of the composite interference, and the optimal measure combination of the composite interference is taken as the final action.

[0011] In an embodiment, the different domains include a frequency domain, a waveform domain, a space domain and a signal processing domain.

[0012] In an embodiment, based on the interference feature vector, feature similarities of the radar transmitting signal and the radar echo signal in different domains are respectively calculated to obtain risk degrees in different domains, including: Based on the interference feature vector, a feature similarity of a carrier frequency of the radar transmitting signal and the radar echo signal in the frequency domain is calculated, and the feature similarity of the carrier frequency is taken as the risk degree in the frequency domain; Based on the interference feature vector, a feature similarity of a waveform correlation coefficient of the radar transmitting signal and the radar echo signal in the waveform domain is calculated, and the feature similarity of the waveform correlation coefficient is taken as the risk degree in the waveform domain; Based on the interference feature vector, a feature similarity of an antenna beam pointing of the radar transmitting signal and the radar echo signal in the space domain is calculated, and the feature similarity of the antenna beam pointing is taken as the risk degree in the space domain; Based on the interference feature vector, a feature similarity of a pulse amplitude of the radar transmitting signal and the radar echo signal in the signal processing domain is calculated, and the feature similarity of the pulse amplitude is taken as the risk degree in the signal processing domain.

[0013] In an embodiment, different domains are combined to obtain a plurality of transformation domains; the transformation domains are encoded to obtain a transformation domain combination vector, including: The frequency domain, the waveform domain, the space domain and the signal processing domain are combined to obtain 15 transformation domains; The 15 transformation domains are binary encoded according to 1 to 15 to obtain an 8-dimensional transformation domain combination vector.

[0014] In an embodiment, the electronic protection anti-interference measure set includes 9 electronic protection anti-interference measures, which are respectively: inter-pulse frequency agility, sub-pulse frequency agility, LFM signal based on frequency modulation, phase encoding based signal, frequency agility signal, adaptive beam forming, sidelobe cancellation, constant false alarm detection and pulse accumulation.

[0015] In an embodiment, when the optimal measure combination in the scope is selected, a priori knowledge base is set as a constraint for selection; The setting prior knowledge base comprises: combining 9 electronic protection anti-interference measures to obtain a plurality of initial combinations; screening all initial combinations that cannot be used in cascade as the prior knowledge base.

[0016] The radar adaptive anti-jamming decision method based on deep learning combines different domains as different agents, splices the transform domain combination vector and the interference category vector as an observation state vector of the agent, inputs the multi-agent reinforcement learning algorithm, and combines the deep learning model, which can not only accurately identify the interference state, but also quickly select the optimal anti-jamming transform domain through multi-domain cooperation optimization between the multi-agents, and generate a multi-technology combination anti-jamming strategy. The optimization of the strategy can be rapidly adjusted in a dynamic complex electromagnetic environment, and the dependence on preset parameters and input state precision is reduced, so that the strategy convergence speed and stability are improved, the reaction speed and adaptability of the system are improved, the anti-jamming capability and overall performance of the radar are greatly improved, and the adaptive anti-jamming of the radar system is realized. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 A flowchart of the radar adaptive anti-jamming decision method based on deep learning in one embodiment is shown. Figure 2 An architecture diagram of the Transformer-LSTM network in one embodiment is shown. Figure 3 An architecture diagram of the MADDPG multi-agent reinforcement learning algorithm in one embodiment is shown. Figure 4 An architecture diagram of the CTDE mechanism in one embodiment is shown. Figure 5 An architecture diagram of the radar adaptive anti-jamming decision method based on deep learning in one embodiment is shown. Figure 6 A structure block diagram of the radar adaptive anti-jamming decision device based on deep learning in one embodiment is shown. Figure 7 An internal structure diagram of the computer device in one embodiment is shown.

[0018] Reference signs: 601 first module, 602 second module, 603 third module, 604 fourth module. DETAILED DESCRIPTION

[0019] In order to make the purposes, technical solutions and advantages of the present application clearer, further detailed description will be given below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not used to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0020] In addition, the description such as "first", "second" and the like in the present application is only for the purpose of description and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" can be explicitly or implicitly included at least one of the features. In the description of the present application, the meaning of "multiple groups" is at least two groups, such as two groups, three groups, etc., unless otherwise specifically limited.

[0021] In the present application, unless otherwise specifically defined and limited, the terms "connection", "fixing" and the like should be understood in a broad sense, for example, "fixing" can be fixed connection, or detachable connection, or integral; can be mechanical connection, or electrical connection, or physical connection or wireless communication connection; can be directly connected, or indirectly connected through intermediate medium, or internal communication of two elements or interaction relationship between two elements, unless otherwise specifically limited. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0022] In addition, the technical solutions of various embodiments of the present application can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can realize it, when the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist, nor within the scope of protection claimed by the present application.

[0023] The present application provides a radar adaptive anti-jamming decision method based on deep learning, as shown in the flowchart, in one embodiment, comprising: Figure 1 As shown in the flowchart, in one embodiment, comprising: Step 101, obtaining a plurality of radar echo signals, after extracting the characteristic parameters, classifying label labeling, and using deep learning algorithm to obtain interference category vector and interference characteristic vector.

[0024] Specifically: Obtaining a plurality of radar echo signals, extracting the characteristic parameters of each radar echo signal, normalizing the characteristic parameters extracted for each radar echo signal, and labeling the signals after normalization (i.e. the category of active jamming) with class labels, and all signals labeled with class labels are combined into a data set (the data set can be divided into a training set and a test set in a ratio of 7:3); The dataset (which consists of multiple 5-dimensional vectors) is input into a deep learning network (e.g., a Transformer-LSTM network) to output interference category vectors and interference feature vectors as interference identification results.

[0025] in: Multiple radar echo signals can be obtained using MATLAB simulation software. For example, MATLAB simulation software can be used to generate time-domain and frequency-domain waveforms of interference signals under different characteristic parameter conditions, which can then be used as multiple radar echo signals.

[0026] The characteristic parameters include: carrier frequency, pulse width, pulse repetition interval, pulse amplitude, and instantaneous bandwidth. The carrier frequency is used to characterize the center frequency of the pulse signal, which helps to distinguish between frequency agility and frequency suppression interference. The pulse width reflects the duration of a single radar pulse and is of great reference value for detecting pulse envelope distortion and wide pulse interference. The pulse repetition interval describes the time interval between adjacent pulses and can help to identify pulse jitter interference and pulse replay interference. The pulse amplitude is used to distinguish between constant amplitude interference and amplitude modulation interference. The instantaneous bandwidth characterizes the bandwidth characteristics of the pulse signal in the frequency domain and is particularly crucial for identifying linear frequency modulation interference and frequency sweep interference.

[0027] Normalization includes: linear normalization, which maps feature parameters of different dimensions to the [0,1] interval, helping to eliminate the interference of differences in feature dimensions on network training. The feature vector obtained after normalization is a continuous real number, not a discretization, but rather retains the physical information of continuous values. The normalization process can be expressed as: ; In the formula, Normalized value For the first a In the nth sample b Original values ​​of the feature parameters, This is the minimum value of the feature parameter across all samples. This is the maximum value of the feature parameter across all samples.

[0028] There are four types of interference: noise interference, false target deception interference, dragging deception interference, and composite interference. Noise interference includes three types: frequency sweeping interference, jamming interference, and targeting interference. Frequency sweeping interference covers the radar's operating frequency band by linearly scanning the frequency over time, forming a continuous suppression, which can be suppressed by pulse frequency agility. Jamming interference uses broadband noise to cover multiple frequency bands simultaneously, causing a large-scale decrease in signal-to-noise ratio, which is usually suppressed by spatial filtering (such as adaptive beamforming). Targeting interference uses narrowband noise to precisely act on the radar's center frequency, suppressing it with high power concentration, which can be suppressed by frequency agility or polarization filtering. The false target deception jamming includes two types, namely, dense false target jamming and range-amplitude deception jamming; the dense false target jamming can be identified and suppressed by using phase coding waveform domain technology, by generating a large number of false targets in the range-Doppler plane to saturate the radar data processing link; the range-amplitude deception jamming can be identified by using time-frequency analysis method, by copying and delaying target echoes to cause range ambiguity and lead to target tracking failure; The deception jamming includes two types, namely, range deception jamming and velocity deception jamming; the range deception jamming can be offset by using Doppler domain filtering, by gradually increasing the echo delay to make the radar range tracking loop deviate from the real target; the velocity deception jamming can be offset by using frequency and waveform joint processing, by gradually changing the Doppler frequency shift to make the velocity tracking loop lose lock and the target lost; The composite jamming includes the combination of noise jamming, false target deception jamming and deception jamming; the composite jamming not only suppresses the radar received signal, but also creates a large number of false targets, causing double damage; for this multi-modal jamming, multi-domain cooperative processing in the frequency domain, spatial domain and waveform domain is needed to achieve comprehensive suppression effect.

[0029] The architecture diagram of the Transformer-LSTM network is shown in Figure 2 The 5-dimensional vector input LSTM, and then passes through the encoder layer, the flattening layer, the full connection layer and the output layer, to output the interference category vector and the interference feature vector.

[0030] The interference category vector is a M dimensional binary coding vector (i.e. 0-1 coding vector, each dimension corresponds to a type of interference category, the corresponding dimension is 1 if the interference category exists, and is 0 if it does not exist), which can be specifically expressed as , M wherein, n represents the number of interference categories, for example: t At time t, there are 8 types of interference categories (sweep jamming, blocking jamming, aiming jamming, dense false target jamming, range-amplitude deception jamming, range deception jamming, velocity deception jamming, composite jamming) and the current is the third type (i.e. M n = 8, m t = 3, ), the interference category vector is ; The interference feature vector refers to the vector composed of feature parameters.

[0031] In this step, feature extraction, normalization processing and label annotation are performed, so as to realize high-precision jamming recognition.

[0032] It should be noted that how to extract feature parameters, how to perform label annotation, how to form a data set, and the Transformer-LSTM network all belong to the prior art and will not be described here.

[0033] Step 102, based on the interference feature vector, the feature similarity of the radar transmitting signal and the radar echo signal in different domains is calculated respectively, and the risk degree in different domains is obtained; the domain with the maximum risk degree is selected as the scope of the anti-interference strategy.

[0034] Specifically: Different domains include: frequency domain, waveform domain, space domain, and signal processing domain; Based on the interference feature vector, the feature similarity of the carrier frequency of the radar transmitting signal and the radar echo signal (i.e. the interference signal) in the frequency domain is calculated, and the feature similarity of the carrier frequency is taken as the risk degree of the frequency domain; Based on the interference feature vector, the feature similarity of the waveform correlation coefficient of the radar transmitting signal and the radar echo signal (i.e. the interference signal) in the waveform domain is calculated, and the feature similarity of the waveform correlation coefficient is taken as the risk degree of the waveform domain; Based on the interference feature vector, the feature similarity of the antenna beam pointing of the radar transmitting signal and the radar echo signal (i.e. the interference signal) in the space domain is calculated, and the feature similarity of the antenna beam pointing is taken as the risk degree of the space domain; Based on the interference feature vector, the feature similarity of the pulse amplitude of the radar transmitting signal and the radar echo signal (i.e. the interference signal) in the signal processing domain is calculated, and the feature similarity of the pulse amplitude is taken as the risk degree of the signal processing domain; Compare the risk degree of the frequency domain, the risk degree of the waveform domain, the risk degree of the space domain, and the risk degree of the signal processing domain. When the domain corresponding to the maximum risk degree is only one, take the domain as the scope of the anti-interference strategy; when the domain corresponding to the maximum risk degree is more than two, take the two or more domains as the scope of the anti-interference strategy (this is a joint domain decision process).

[0035] Wherein: The feature similarity is measured by the distance between the sample features, which can be calculated by using the Euclidean distance divided by the number of domain features: ; In the formula, is the sample is the sample The feature similarity of the first domain, the smaller the value, the higher the similarity; is the first feature number; is the total number of features selected in a domain; is the sample In the first characteristics in the first characteristics in the first characteristics selected for the sample in the first In the first characteristics in the first characteristics in the first characteristics selected for the sample in the first characteristics selected for the sample in the first

[0036] In this step, the frequency domain selects the carrier frequency as the feature, the waveform domain selects the correlation coefficient (Corr) as the feature, the spatial domain selects the antenna beam pointing as the feature, and the signal processing domain selects the pulse amplitude as the feature. Not only can interference be applied in a certain domain (frequency domain, waveform domain, spatial domain, signal processing domain), but also the working parameters of the radar in these domains are "aligned" with the interference signal, that is, the parameters of the interference signal and the radar in the domain are matched with each other, thereby improving the better interference effect of the interference signal on the radar.

[0037] On this basis, the feature similarity is calculated to quantify the risk degree of the current interference in each domain, reflect the urgency of the interference measure requirement in the domain, and if the feature similarity is smaller and the domain of the maximum feature similarity is more (the number of accurately matched domains is more), the risk degree of the interference is greater, the domain alignment degree is higher, the effective energy of the interference signal at the radar end is stronger, the radar performance decreases more obviously, and the requirement for anti-jamming action in this domain is greater, thereby determining the most targeted defense domain and realizing domain optimization. At the same time, by determining the anti-jamming in which domain is most effective, the selection range of the subsequent selection from the electronic protection anti-jamming measure set (all ECCM technologies) can be reduced, thereby reducing the algorithm calculation amount, speeding up the decision-making speed, and reducing the risk of selecting the wrong electronic protection anti-jamming measure.

[0038] It should be noted that the frequency domain, the waveform domain, the spatial domain, and the signal processing domain all belong to the prior art, and will not be described here.

[0039] In step 103, different domains are taken as different agents; different domains are combined to obtain a plurality of transformation domains; the transformation domains are encoded to obtain a transformation domain combination vector; and the transformation domain combination vector and the interference category vector are spliced to obtain an observation state vector of the agent.

[0040] Specifically: The frequency domain, the waveform domain, the spatial domain, and the signal processing domain are taken as different agents, and there are a total of 4 agents, i.e. ; The frequency domain, the waveform domain, the spatial domain, and the signal processing domain are combined to obtain 15 transformation domains, i.e. ; The 15 transform domains are encoded in binary according to the numbers 1 to 15 to obtain a combined transform domain vector. The observed state vector of the agent is obtained by concatenating the transformed domain combination vector and the interference category vector.

[0041] in: The transform-domain combined vector is a 4-dimensional binary encoded vector (i.e., a 0-1 encoded vector), for example: t At time 15, there are 15 transform domains, and the current one is the 3rd type (i.e., v When =3), the transform domain combination vector is .

[0042] The agent's observation state vector is a 12-dimensional binary encoded vector (i.e., a 0-1 encoded vector), obtained by concatenating a transform-domain combined vector and an interference category vector. For example, t At time t, the transform domain combined vector is The interference category vector is , No. n The observation state vector of each agent is: .

[0043] In this step, the transform domain combination vector and the interference category vector are concatenated to obtain the agent's observation state vector, which is then used as the input to the multi-agent reinforcement learning algorithm.

[0044] Step 104: Using the agent's observation state vector as the input to the multi-agent reinforcement learning algorithm, the agent selects the optimal combination of measures in the scope of action from the set of electronic protection and anti-jamming measures as the final action, and outputs it by the multi-agent reinforcement learning algorithm to realize radar adaptive anti-jamming decision.

[0045] Specifically: The observed state vector of the agent is used as the input to the multi-agent reinforcement learning algorithm; The set of electronic protection and anti-interference measures includes nine types of electronic protection and anti-interference measures, namely: inter-pulse frequency agility, sub-pulse frequency agility, frequency modulation-based LFM signal, phase-coded signal, frequency agility signal, adaptive beamforming, sidelobe cancellation, constant false alarm rate detection, and pulse accumulation. The agent selects a combination of measures within the scope of the electronic protection and anti-interference measures set, and sets a prior knowledge base as a constraint for the selection to obtain the optimal combination of measures, and uses the optimal combination of measures as the final action; The final action is output by a multi-agent reinforcement learning algorithm to achieve adaptive anti-jamming decision-making for the radar.

[0046] More specifically: The observed state vector of the agent is used as the input to the MADDPG multi-agent reinforcement learning algorithm; The set of electronic protection and anti-interference measures includes nine electronic protection and anti-interference measures (ECCM): inter-pulse frequency agility, sub-pulse frequency agility, frequency modulation-based LFM signal, phase-coded signal, frequency agility signal, adaptive beamforming, sidelobe cancellation, constant false alarm rate detection, and pulse accumulation. When the interference category is noise interference, the agent selects a combination of measures within its scope from the set of electronic protection and anti-interference measures. A prior knowledge base is set up (combining the nine electronic protection and anti-interference measures to obtain multiple initial combinations; filtering out all initial combinations that cannot be cascaded, which serve as the prior knowledge base), acting as a constraint on the selection (if the selected measure combination exists in the prior knowledge base, the action is deemed invalid, and a new measure combination is selected until the current measure combination does not exist in the prior knowledge base, at which point the action is deemed valid, ensuring the rationality of the output action), resulting in a valid combination; the difference in SINR before and after adopting the valid combination is used as the reward function. The process is iterated until the SINR difference before and after taking a legal combination meets the preset first condition. The current legal combination is the optimal combination of noise interference measures, and the optimal combination of noise interference measures is taken as the final action. When the interference category is false target deception interference, the agent selects a combination of measures within its scope from the set of electronic protection and anti-interference measures. A prior knowledge base is set up (combining nine electronic protection and anti-interference measures to obtain multiple initial combinations; filtering out all initial combinations that cannot be cascaded, serving as the prior knowledge base), which acts as a constraint on the selection (if the selected measure combination exists in the prior knowledge base, the action is deemed invalid, and a new measure combination is selected until the current measure combination does not exist in the prior knowledge base, at which point the action is deemed valid, ensuring the rationality of the output action), resulting in a valid combination; the real target recognition rate (i.e., the ratio of the number of real targets actually detected by the radar to the total number of targets detected by the system) is used as the reward function. The process is iterated until the real target recognition rate meets the preset second condition. The current combination of measures is the optimal combination of measures for deceiving and interfering with false targets. The optimal combination of measures for deceiving and interfering with false targets is used as the final action. When the interference category is the towed deception jamming, the agent selects the scope measure combination in the electronic protection anti-jamming measure set, sets the priori knowledge base (combines the 9 electronic protection anti-jamming measures to obtain a plurality of initial combinations; screens all initial combinations that cannot be used in cascade as the priori knowledge base), as the constraint of selection (if the selected measure combination exists in the priori knowledge base, it is judged that the action is illegal, the measure combination is reselected, and until the current measure combination does not exist in the priori knowledge base, it is judged that the action is legal, so as to ensure the rationality of the output action), to obtain a legal combination; the negative value of the tracking accuracy error is taken as a reward function ( ), and the loop iteration is performed until the negative value of the tracking accuracy error meets a preset third condition, and the current measure combination is the optimal measure combination of the towed deception jamming, and the optimal measure combination of the towed deception jamming is taken as the final action; When the interference category is the composite jamming, the agent selects the scope measure combination in the electronic protection anti-jamming measure set, sets the priori knowledge base (combines the 9 electronic protection anti-jamming measures to obtain a plurality of initial combinations; screens all initial combinations that cannot be used in cascade as the priori knowledge base), as the constraint of selection (if the selected measure combination exists in the priori knowledge base, it is judged that the action is illegal, the measure combination is reselected, and until the current measure combination does not exist in the priori knowledge base, it is judged that the action is legal, so as to ensure the rationality of the output action), to obtain a legal combination; the weighted sum of the difference between the SINR before and after the measure combination is taken, the real target recognition rate and the negative value of the tracking accuracy error (how to take each weight value is the prior art, and specific can be flexibly adjusted according to the interference type contained in the composite jamming, and details are not described here) is taken as a reward function ( ), and the loop iteration is performed until the weighted sum meets a preset fourth condition, and the current measure combination is the optimal measure combination of the composite jamming, and the optimal measure combination of the composite jamming is taken as the final action; The final action is output by the MADDPG multi-agent reinforcement learning algorithm and is executed by the radar system, so as to realize the radar adaptive anti-jamming decision.

[0047] Among them: The electronic protection anti-jamming measures contained in each domain are shown in Table 1: Table 1: Electronic protection anti-jamming measures

[0048] The final action (i.e. the optimal measure combination of the interference) output by the MADDPG multi-agent reinforcement learning algorithm is a W dimension binary coding vector (i.e. 0-1 coding vector, each dimension corresponds to an electronic protection anti-jamming measure, and the corresponding dimension is 1 when the electronic protection anti-jamming measure is selected, and is 0 when the electronic protection anti-jamming measure is not selected), and can be specifically expressed as , W This indicates the number of electronic protection and anti-interference measures selected within the scope of action, for example: t At any given time, there are four electronic protection and anti-interference measures within the scope, and the second one (i.e.) is selected. W =4、 When ), the final action is .

[0049] In this step, within the selected domain, action loop optimization is performed by combining multi-agent reinforcement learning algorithms and a set of electronic protection and anti-jamming measures to achieve radar anti-jamming decision-making.

[0050] like Figure 3 The diagram illustrates the architecture of the MADDPG multi-agent reinforcement learning algorithm. The overall system consists of three parts: the environment, multiple agents (taking i agents as an example; the diagram only shows agent 1, agent 2, agent 3, and agent i; all agents have the same structure, employing a dual-network design of "online network + target network"), and an experience replay pool, forming a closed-loop interaction. The complex electromagnetic environment in which the radar operates outputs the corresponding state perceived by each agent. This refers to the fusion representation of current environmental information, specifically a 12-dimensional observation state vector composed of an 8-dimensional interference category vector and a 4-dimensional transform domain combination vector, reflecting "current interference type + domain to be defended". Each agent then outputs an action based on the state. This involves selecting a binary-coded combination vector from nine electronic protection and anti-interference measures to represent the anti-interference operation to be performed in that domain, such as inter-pulse frequency agility and adaptive beamforming. The first three agents in the diagram follow the same process: the online policy network determines the anti-interference operation based on the current state... Real-time output of actions (in (For online policy network parameters), the target policy network then determines the next state. Real-time output of actions (in (These are the network parameters for the target policy), used for subsequent target value calculation; the corresponding online value network evaluates the online value of the state-action pair. (in (For the online value network parameters), the target value network provides a stable value estimate and outputs the target value. Agent i further refines the training process: the online policy network outputs... Then, the superimposed OU noise forms the final action. To balance exploration and utilization; online value output by online value networks The target value output by the target value network (in For the target value network parameters, Output for other intelligent agents The summation of the two values ​​(the summation of the ... GB Update online network parameters Then, soft updates are used to slowly synchronize the online network parameters to the target policy network parameters and the target value network parameters to maintain training stability. The experience replay pool is used to store experience tuples generated by the agent's interactions with the environment. ,in, s This is the current state. s ′ represents the next state. r The reward for environmental feedback (such as the improvement in SINR after interference immunity) is input into the agent as... r i ), a Actions performed by the intelligent agent d The flag indicates whether a round has ended; during training, small batches of samples are randomly sampled for network updates to break data correlation and improve learning efficiency. All parameters in the figure collectively embody the collaborative decision-making and closed-loop optimization mechanism of multi-agent systems in various radar anti-jamming domains. Stable training is achieved through a dual-network structure and experience replay, ultimately outputting a combination of anti-jamming measures adapted to complex electromagnetic environments.

[0051] Because different anti-jamming technologies (electronic protection anti-jamming measures) correspond to different signal processing links and physical principles, their effects may reinforce or cancel each other out, or fail under specific environments. Therefore, it is difficult for a single agent to simultaneously optimize the use strategy of multiple technologies. The MADDPG multi-agent reinforcement learning algorithm enables each agent to focus on a single domain and, during training, utilizes a centralized training and distributed execution (CTDE) mechanism (such as... Figure 4 As shown, this illustrates the collaborative relationship between multiple agent networks (Actor network) and a centralized Critic network during the training phase to achieve policy coordination.

[0052] A priori knowledge base is set up (combining 9 electronic protection and anti-interference measures to obtain multiple initial combinations; filtering out all initial combinations that cannot be cascaded as the priori knowledge base) as a constraint for selection (if the selected measure combination exists in the priori knowledge base, the action is judged as invalid, and a new measure combination is selected until the current measure combination does not exist in the priori knowledge base, then the action is judged as valid, so as to ensure the rationality of the output action), resulting in valid combinations. This can avoid the problem that electronic protection and anti-interference measures in some domains cannot be used simultaneously during the anti-interference process, and avoid the failure of anti-interference effect or radar system failure due to ECCM compatibility conflicts; for example, the electronic protection and anti-interference measures for pulse agility waveforms in the frequency domain and linear frequency modulation (LFM) signals in the waveform domain cannot be used simultaneously.

[0053] For noise interference, its core mechanism is to reduce the ratio of "useful signal" to "interference + background noise" received by the detection system by superimposing random noise, ultimately causing the useful signal to be submerged. The signal-to-interference-plus-noise ratio (SINR) can directly quantify the degree of damage to "signal quality" caused by noise interference. This application uses the difference in SINR before and after adopting a legal combination as the reward function, so as to most intuitively reflect whether the anti-interference strategy effectively counteracts and suppresses the "signal submersion" problem, and judge the effectiveness of the radar anti-interference strategy.

[0054] The core mechanism of deceptive jamming against false targets is to generate a large number of false target signals by releasing radar decoys and forging false echoes. These signals interfere with the target recognition process of the detection system, making it difficult to distinguish between "real targets" and "false targets". Ultimately, this leads to the system mistracking false targets and missing real targets. This application uses the real target recognition rate (i.e., the ratio of the number of real targets actually detected by the radar to the total number of targets detected by the system) as the reward function. This can directly reflect the detection system's ability to distinguish real targets in a mixed scenario of real and false targets, quantify the anti-jamming effect of this type of jamming, and thus accurately judge the effectiveness of the anti-jamming strategy.

[0055] The core mechanism of deception interference is to release progressive false signals to trick the tracking module of the detection system into deviating, causing a deviation in the tracking of the real target. This application uses the negative value of the tracking accuracy error as a reward function to quantify the deviation between the system's tracking position and the actual position of the real target. The smaller the error, the more effective the anti-interference. This indicator directly corresponds to the "tracking deviation" essence of interference and can reflect the system's ability to maintain the tracking accuracy of the real target. After negative reward design, it can be transformed into a reward value that conforms to the logic of "the better the anti-interference, the higher the reward".

[0056] For complex interference, its core mechanism is to damage the system from multiple dimensions such as signal, recognition, and tracking. This application uses the weighted sum of the SINR difference before and after taking measures, the true target recognition rate, and the negative value of the tracking accuracy error as the reward function. It can reflect the priority difference of different tasks to the performance of each dimension through weight setting, comprehensively quantify the overall anti-interference capability of the system under the simultaneous action of multiple interferences, avoid the one-sidedness of a single indicator, and comprehensively evaluate the anti-interference effect of such multi-dimensional damage.

[0057] It should be noted that the MADDPG multi-agent reinforcement learning algorithm, electronic protection and anti-jamming measures, how to select initial combinations that cannot be cascaded, how to calculate the SINR difference, how to obtain the actual number of targets detected by the radar, how to obtain the total number of targets detected by the system, how to obtain the tracking accuracy error, the first condition, the second condition, the third condition, and the fourth condition are all existing technologies and will not be elaborated here.

[0058] In this application, as Figure 5 As shown, firstly, interference is identified and 5-dimensional features are extracted, outputting interference category vectors and interference feature vectors. Next, domain optimization is performed, calculating the similarity of each domain to determine the optimal scope. Subsequently, combined encoding is performed, treating each domain as an agent and obtaining the agent's observation state vector. Finally, policy output is performed, employing a multi-agent reinforcement learning framework to collaboratively optimize the combination of anti-interference techniques within the scope, introducing a prior knowledge base to ensure policy compatibility, and constructing a reward function through multiple indicators to achieve closed-loop policy optimization.

[0059] The aforementioned deep learning-based radar adaptive anti-jamming decision-making method treats different domains as different agents. It concatenates the transform domain combination vector and the interference category vector to form the agent's observation state vector, which is then input into a multi-agent reinforcement learning algorithm. Combined with a deep learning model, this method can accurately identify interference states and quickly select the optimal anti-jamming transform domain through multi-domain collaborative optimization among multiple agents. It also generates a multi-technology combined anti-jamming strategy. The strategy optimization can be rapidly adjusted in dynamic and complex electromagnetic environments, reducing dependence on preset parameters and input state accuracy. This improves strategy convergence speed and stability, enhances system response speed and adaptability, significantly improves radar anti-jamming capability and overall performance, and achieves adaptive anti-jamming for the radar system.

[0060] Specifically, this application has the following beneficial effects: 1) Unlike traditional methods that rely solely on simple energy detection or shallow classifiers, this application maps the core domain of radar anti-jamming to independent agents. Each agent focuses on optimizing anti-jamming measures within its own domain. Simultaneously, it transforms the actual features of the transformation domain combination and jamming category into vector form, and uses the concatenated vectors to construct the observation state space. This vector space serves as the input to a multi-agent reinforcement learning algorithm, achieving a fusion representation of multi-domain state information. This overcomes the limitations of traditional methods that rely on pre-defined parameter strategies. Furthermore, by combining this with a deep learning model (which is an existing technology), it can effectively address the limitations of radar anti-jamming by optimizing the anti-jamming measures within its own domain. It accurately captures, automatically extracts, and effectively processes complex spatiotemporal features from radar echo signals. It can not only effectively process high-dimensional feature signals of long-term sequences, taking into account both local time dependence and global pattern extraction, but also specifically address scenarios where multiple interference signals act simultaneously. Even in complex electromagnetic environments with low signal-to-noise ratios and overlapping types of interference, it can still maintain a high interference identification accuracy, achieving high-precision interference identification and pattern classification. This significantly improves the system's robustness, reliability, and adaptive anti-interference capability, making the radar more stable in complex electromagnetic environments. It can quickly adapt to environmental changes and effectively cope with variable interference, ensuring the efficient operation of the radar system.

[0061] 2) Compared with traditional rule-driven methods, the features in this application work together to overcome the shortcomings of existing technologies where agent reinforcement learning can only optimize single-dimensional anti-interference strategies. It allows agents from different domains to work collaboratively through centralized training and decentralized execution, which not only solves the compatibility conflict problem between multi-domain measures, but also narrows the decision space through the strategy of "prioritizing the domain of risk level", greatly improving the convergence speed and decision accuracy of the algorithm. At the same time, this application can provide richer feature representation and correlation information, thereby more accurately capturing the multi-dimensional dynamic characteristics of interference signals, improving the convergence speed and flexibility of interference strategies, and directly further promoting the significant improvement of system robustness, reliability and adaptive anti-interference capability, ensuring interference effect, reducing human intervention, improving the survivability and stability of detection and tracking, and enabling the radar system to quickly adapt to environmental changes and effectively cope with variable interference, ensuring stable and efficient operation in complex electromagnetic environments.

[0062] This application is applicable to the field of autonomous decision-making technology for radar anti-jamming, especially radar anti-jamming scenarios in complex electromagnetic environments (such as radar anti-jamming systems), and has broad application prospects.

[0063] It should be understood that, although Figure 1The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0064] This application also provides a radar adaptive anti-jamming decision-making device based on deep learning, such as... Figure 6 As shown, in one embodiment, it includes: a first module 601, a second module 602, a third module 603, and a fourth module 604, wherein: The first module 601 is used to acquire multiple radar echo signals, extract feature parameters, label them with category labels, and use deep learning algorithms to obtain interference category vectors and interference feature vectors (i.e., interference identification module). The second module 602 is used to calculate the feature similarity between radar transmitted signals and radar echo signals in different domains based on interference feature vectors, and obtain the risk level of different domains; select all domains with the highest risk level as the domain of application of the anti-interference strategy (i.e., domain optimization module). The third module 603 is used to treat different domains as different agents; combine different domains to obtain multiple transform domains; encode the transform domains to obtain a transform domain combination vector; and concatenate the transform domain combination vector and the interference category vector to obtain the agent's observation state vector (i.e., the encoding module). The fourth module 604 is used to take the observation state vector of the agent as the input of the multi-agent reinforcement learning algorithm. The agent selects the optimal combination of measures in the scope of the electronic protection and anti-jamming measures set, and outputs it as an action by the multi-agent reinforcement learning algorithm to realize radar adaptive anti-jamming decision (i.e., ECCM selection module).

[0065] Specific limitations regarding the deep learning-based radar adaptive anti-jamming decision-making device can be found in the limitations of the deep learning-based radar adaptive anti-jamming decision-making method described above, and will not be repeated here. Each module in the aforementioned device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in a computer device, or stored in software in the memory of a computer device, so that the processor can call and execute the operations corresponding to each module.

[0066] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a deep learning-based radar adaptive anti-jamming decision-making method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0067] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0068] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the method described above.

[0069] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.

[0070] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0071] The contents not described in detail in this specification are existing technologies known to those skilled in the art.

[0072] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0073] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended application documents.

Claims

1. A radar adaptive anti-jamming decision-making method based on deep learning, characterized in that, include: Multiple radar echo signals were acquired, and after extracting feature parameters, category labels were applied. Then, a deep learning algorithm was used to obtain the interference category vector and the interference feature vector. Based on the interference feature vector, the feature similarity between the radar transmitted signal and the radar echo signal in different domains is calculated to obtain the risk level in different domains; the domain with the highest risk level is selected as the domain of application of the anti-jamming strategy. Different domains are treated as different agents; different domains are combined to obtain multiple transform domains; the transform domains are encoded to obtain a transform domain combination vector; the transform domain combination vector and the interference category vector are concatenated to obtain the agent's observation state vector; Using the observation state vector of the agent as the input of the multi-agent reinforcement learning algorithm, the agent selects the optimal combination of measures in the scope of the electronic protection and anti-jamming measures set as the final action, and the multi-agent reinforcement learning algorithm outputs the result to realize radar adaptive anti-jamming decision.

2. The radar adaptive anti-jamming decision-making method based on deep learning according to claim 1, characterized in that, Multiple radar echo signals were acquired, and after extracting feature parameters, category labels were applied. A deep learning algorithm was then used to obtain the interference category vector and interference feature vector, including: Multiple radar echo signals are acquired, feature parameters of each radar echo signal are extracted, the extracted feature parameters are normalized, and the normalized signals are labeled with category labels. All the labeled signals are combined into a dataset. The dataset is input into a deep learning network to output a disturbance category vector and a disturbance feature vector.

3. The radar adaptive anti-jamming decision-making method based on deep learning according to claim 2, characterized in that, Characteristic parameters include: carrier frequency, pulse width, pulse repetition interval, pulse amplitude, and instantaneous bandwidth; The carrier frequency is used to characterize the center frequency of the pulse signal; the pulse width reflects the duration of a single radar pulse; the pulse repetition interval describes the time interval between adjacent pulses; the pulse amplitude is used to distinguish between constant amplitude interference and amplitude modulation interference; and the instantaneous bandwidth characterizes the bandwidth characteristics of the pulse signal in the frequency domain.

4. The radar adaptive anti-jamming decision-making method based on deep learning according to claim 3, characterized in that, The types of interference include: noise interference, decoy target deception interference, dragging deception interference, and composite interference; Noise interference includes: frequency sweeping interference, jamming interference, and aiming interference; false target deception interference includes: dense false target interference and range-amplitude deception interference; drag deception interference includes: range drag interference and velocity drag interference; composite interference includes: a combination of noise interference, false target deception interference, and drag deception interference.

5. The radar adaptive anti-jamming decision-making method based on deep learning according to claim 4, characterized in that, The agent selects the optimal combination of measures within its scope from the set of electronic protection and anti-interference measures as its final action, including: When the interference category is noise interference, the agent selects a combination of measures within the scope of the electronic protection anti-interference measures set, uses the SINR difference before and after taking the measure combination as the reward function, and performs iterative loops until the SINR difference before and after taking the measure combination meets the preset first condition. The current measure combination is the optimal measure combination for noise interference, and the optimal measure combination for noise interference is taken as the final action. When the interference category is false target deception interference, the agent selects the combination of measures within the scope of the electronic protection anti-interference measures set, uses the real target recognition rate as the reward function, and performs iterative loops until the real target recognition rate meets the preset second condition. The current combination of measures is the optimal combination of measures for false target deception interference, and the optimal combination of measures for false target deception interference is used as the final action. When the interference category is dragging deception interference, the agent selects a combination of measures within the scope of the electronic protection anti-interference measures set, uses the negative value of the tracking accuracy error as the reward function, and performs iterative loops until the negative value of the tracking accuracy error meets the preset third condition. The current combination of measures is the optimal combination of measures for dragging deception interference, and the optimal combination of measures for dragging deception interference is used as the final action. When the interference category is complex interference, the agent selects a combination of measures within the scope of the electronic protection anti-interference measures set. The reward function is the weighted sum of the SINR difference before and after taking the measure combination, the true target recognition rate, and the negative value of the tracking accuracy error. The process is iterated until the weighted sum meets the preset fourth condition. The current measure combination is the optimal measure combination for complex interference, and the optimal measure combination for complex interference is taken as the final action.

6. The radar adaptive anti-jamming decision-making method based on deep learning according to any one of claims 1 to 5, characterized in that, Different domains include: frequency domain, waveform domain, spatial domain, and signal processing domain.

7. The radar adaptive anti-jamming decision-making method based on deep learning according to claim 6, characterized in that, Based on the interference feature vector, the feature similarity between the radar transmitted signal and the radar echo signal in different domains is calculated to obtain the risk level in different domains, including: Based on the interference feature vector, the feature similarity between the radar transmitted signal and the radar echo signal in the frequency domain is calculated, and the feature similarity of the carrier frequency is used as the risk level in the frequency domain. Based on the interference feature vector, the feature similarity of the waveform correlation coefficient between the radar transmitted signal and the radar echo signal in the waveform domain is calculated, and the feature similarity of the waveform correlation coefficient is used as the risk level in the waveform domain. Based on the interference feature vector, the feature similarity of the antenna beam pointing of the radar transmitted signal and the radar echo signal in the airspace is calculated, and the feature similarity of the antenna beam pointing is used as the risk level in the airspace. Based on the interference feature vector, the characteristic similarity of the pulse amplitude of the radar transmitted signal and the radar echo signal in the signal processing domain is calculated, and the characteristic similarity of the pulse amplitude is used as the risk level in the signal processing domain.

8. The radar adaptive anti-jamming decision-making method based on deep learning according to any one of claims 1 to 5, characterized in that, By combining different domains, multiple transformation domains can be obtained; Encoding the transform domain yields a combined transform domain vector, including: By combining the frequency domain, waveform domain, spatial domain, and signal processing domain, 15 transform domains are obtained; The 15 transform domains are encoded in binary according to the numbers 1 to 15 to obtain an 8-dimensional transform domain combination vector.

9. The radar adaptive anti-jamming decision-making method based on deep learning according to any one of claims 1 to 5, characterized in that, The set of electronic protection and anti-interference measures includes nine types of electronic protection and anti-interference measures, namely: inter-pulse frequency agility, sub-pulse frequency agility, frequency modulation-based LFM signal, phase-coded signal, frequency agility signal, adaptive beamforming, sidelobe cancellation, constant false alarm rate detection, and pulse accumulation.

10. The radar adaptive anti-jamming decision-making method based on deep learning according to claim 9, characterized in that, When selecting the optimal combination of measures for a given scope, a prior knowledge base is set up as a constraint for the selection. Setting up the prior knowledge base includes: combining nine electronic protection and anti-interference measures to obtain multiple initial combinations; and filtering out all initial combinations that cannot be cascaded as the prior knowledge base.

Citation Information

Patent Citations

  • Multi-unmanned aerial vehicle action decision-making method and device based on reinforcement learning

    CN111708355A

  • Radar anti-interference intelligent decision-making method based on reinforcement learning

    CN113625233A

  • Radar intelligent anti-interference decision-making method based on multi-criterion multi-cost function

    CN117148286A

  • Radar anti-interference decision-making method based on recursive game training reinforcement learning

    CN120195629A

  • Radar deception jamming identification method based on adversarial game optimization

    CN120254770A