Radar deception jamming identification method based on adversarial game optimization
By building an adaptive risk game model of radar system and deceiving interference and combining deep learning and reinforcement learning methods, the problem of deceiving interference recognition of radar system in complex electromagnetic environments is solved, efficient and adaptive recognition and classification are achieved, and the anti-interference ability and recognition accuracy of the radar system are improved.
Patent Information
- Application Number
- CN202510370617.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When existing radar systems face deception interference in complex electromagnetic environments, they have low recognition accuracy and poor adaptability, making it difficult to achieve effective real-time adaptation and optimization, and traditional methods are difficult to cope with intelligent and dynamic deception interference.
Adopting an adversarial game optimization method is adopted to build an adaptive risk game model between the radar system and the deceptive interference, combining the deep learning identification network and reinforcement learning module, dynamic identification and classification of the radar system is realized through adaptive risk control and strategy optimization.
It improves the anti-interference ability of the radar system in complex electromagnetic environments, maintains high recognition accuracy, reduces false alarm rates and missed alarm rates, and can adaptively adjust the identification strategy in a dynamic environment to reduce the uncertainty of combat decisions.
Smart Images

Figure CN120254770A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of radar, and in particular to a method for identifying radar deception jamming based on adversarial game optimization. Background Technique
[0002] With the continuous development of radar detection technology, radar systems are widely used in target detection, tracking and recognition in the fields of military, aerospace and autonomous driving. However, in the face of the increasingly complex electromagnetic environment and confrontation conditions, the detection ability of radar systems has been severely challenged by deception jamming technology. Deception jamming is an adversarial means that makes the radar system misjudge by simulating real target echo signals. Common deception jamming methods include delay deception, repeater deception, and false target jamming. These deception means can interfere with the normal operation of the radar system, making it difficult to accurately distinguish real targets from interference signals, thus affecting the reliability of tactical decisions.
[0003] In the prior art, the identification of radar deception jamming mainly relies on traditional signal processing methods, such as matched filtering, space-time adaptive processing (STAP), and polarization analysis. Traditional methods can enhance the anti-jamming ability of radar systems to a certain extent, but there are the following problems and limitations in practical applications: First, traditional signal processing methods perform limitedly in the face of complex deception jamming. Traditional methods usually rely on prior knowledge or fixed signal features for classification, and when facing new or unknown types of deception jamming, the identification effect is poor. In addition, due to the increasingly intelligent and dynamic means of deception jamming, it is difficult for traditional methods to achieve effective real-time adaptation in complex environments. Second, some radar systems have introduced deception jamming detection methods based on pattern recognition and machine learning. For example, support vector machines, random forests, and deep neural networks have certain advantages in complex environments, but there are still the following problems: On the one hand, machine learning methods usually rely on a large amount of labeled data for training, and it is difficult to obtain deception jamming data, resulting in insufficient training data; on the other hand, the generalization ability of traditional machine learning models is limited, and it is difficult to adapt to variable deception jamming strategies, and the reliability in actual combat environments is low.
[0004] In addition, although existing deep learning methods can automatically extract the features of deception jamming, they still face many challenges in practical applications. On the one hand, since deep learning models usually adopt static training methods, it is difficult to make adaptive adjustments in the face of dynamically changing deception jamming strategies; on the other hand, existing deep learning methods lack an optimization strategy combined with game theory and cannot effectively model the adversarial relationship between the radar system and deception jamming, resulting in the inability to fully improve the identification accuracy and anti-jamming ability.
[0005] In summary, there are still problems in the existing technology for radar deception jamming recognition, such as low recognition accuracy, poor adaptability, and insufficient anti-interference ability, which are difficult to meet the requirements of efficient and intelligent anti-jamming in complex electromagnetic environments. Therefore, there is an urgent need for a new radar deception jamming recognition method that can combine deep learning and game optimization to achieve adaptive recognition and optimization in an adversarial environment and improve the stability and reliability of radar systems in deception jamming environments. Summary of the Invention
[0006] An object of the present invention is to propose a radar deception jamming recognition method based on adversarial game optimization, which enables the radar system to effectively avoid high-risk strategies, thereby reducing the uncertainty in combat decision-making.
[0007] A radar deception jamming recognition method based on adversarial game optimization according to an embodiment of the present invention includes the following steps:
[0008] S1. Obtain radar echo signal data, including real target echo signals and deception jamming signals;
[0009] S2. Perform noise suppression, clutter suppression, and time-domain and frequency-domain feature extraction on the radar echo signal data to obtain preprocessed radar echo signal data;
[0010] S3. Construct an adaptive risk game model between the radar and the deception jamming based on the preprocessed radar echo signal data;
[0011] S4. Use the preprocessed radar echo signal data to train a deep learning recognition network, which is used to extract the hidden features of deception jamming signals and achieve a preliminary classification of deception jamming signals and real target echo signals;
[0012] S5. Introduce a reinforcement learning module to integrate the deep learning recognition network and the adaptive risk game model. The reinforcement learning module obtains a reward signal through interaction with the combat environment and adaptively adjusts the radar recognition strategy and the parameters of the deep learning recognition network according to the reward signal;
[0013] S6. Apply the optimized radar recognition strategy to the processing of real-time radar echo signal data, recognize and classify deception jamming signals in real time, and output corresponding recognition results.
[0014] Optionally, the S2 specifically includes:
[0015] S21. Sample the radar echo signal data x(t), where the radar echo signal data x(t) consists of a real target echo signal s(t), a deception jamming signal j(t), and noise n(t);
[0016] S22. Perform noise suppression, adopt an adaptive filtering method to reduce Gaussian white noise interference, and obtain the filtered radar echo signal data;
[0017] S23. Perform clutter suppression, enhance the real target signal component, reduce ground clutter and sea clutter interference, and obtain the radar echo signal data after clutter suppression;
[0018] S24. Extract time-domain and frequency-domain features, including short-time energy, instantaneous frequency, and time-frequency joint distribution, and obtain the preprocessed radar echo signal data X processed (n).
[0019] Optionally, the specific steps of S3 are as follows:
[0020] S31. Construct an adaptive risk game model for the radar system and deceptive jamming. Assume that the radar system is the game player R and the deceptive jamming is the opponent J. Among them, the radar system R distinguishes the real target echo signal from the deceptive jamming signal through the optimal recognition strategy, and the deceptive jamming J confuses the radar system through the jamming strategy to make it misidentify. Obtain the preprocessed radar echo signal data X processed (n) as the basis for the radar system to make a decision on the recognition strategy in the game. Combining the time-domain and frequency-domain features in the radar echo signal data, define the recognition strategy set of the radar system as S R , and define the jamming strategy set of the deceptive jamming as S J :
[0021]
[0022] Among them, represents the i-th recognition strategy of the radar system, represents the j-th jamming strategy of the deceptive jamming, and m and n are the numbers of strategies of the radar system and the deceptive jamming respectively;
[0023] S32. Establish a profit function based on the preprocessed radar echo signal data. Define the profit function of the radar system and the profit function of the deceptive jamming The two satisfy the zero-sum game relationship:
[0024]
[0025] Among them, the profit function of the radar system is driven by the feature data extracted from the preprocessed radar echo signal data and depends on the recognition accuracy rate P acc , the false alarm rate P fa and the missed alarm rate P md , and the profit function of the deceptive jamming reflects its ability to successfully deceive the radar system:
[0026]
[0027] Among them, w1, w2, and w3 are the weight coefficients of the recognition accuracy rate, false alarm rate, and missed alarm rate respectively;
[0028] S33. Introduce an adaptive risk control mechanism, calculate the risk return of the radar system under different strategy combinations based on the radar echo signal data, calculate the maximum loss borne by the radar system using conditional value at risk, optimize the recognition strategy and reduce false alarms and missed alarms of risk situations above the risk threshold. If the risk threshold of the radar system under conditional value at risk is α, the risk return is calculated as follows:
[0029]
[0030] Among them, represents the expected return of the radar system;
[0031] S34. Optimize the game strategy and solve the Nash equilibrium, and use the game optimization method to calculate the optimal mixed recognition strategy of the radar system and the optimal mixed interference strategy of deception interference Solve the mixed strategy Nash equilibrium point that satisfies the following relationship:
[0032]
[0033] S35. Apply the calculated optimal recognition strategy to real-time radar signal processing, combine the radar echo signal data X processed (n) to identify and classify deception interference signals in real time, output the final recognition result, and through the risk adaptive adjustment mechanism, if the risk return is lower than the set threshold α during the recognition process, the radar system automatically adjusts the recognition strategy.
[0034] Optionally, the recognition accuracy rate is the probability that the radar successfully distinguishes the real echo signal from the deception interference signal, the false alarm rate is the probability that the radar misidentifies the deception interference as a real target, and the missed alarm rate is the probability that the radar misidentifies a real target as a deception interference.
[0035] Optionally, the specific content of S4 includes:
[0036] S41. Use the preprocessed radar echo signal data X processed (n) as the input of the deep learning recognition network D(θ), where θ is the network parameter;
[0037] S42. Design a deep learning recognition network D(θ). The deep learning recognition network includes a feature extraction layer and a classifier layer. The feature extraction layer is used to extract the hidden features of the deception interference signal from the preprocessed radar echo signal data and output a feature vector f(n);
[0038] S43. Use the classifier layer to map the extracted feature vector f(n) to a class probability where The value range of is [0,1], which represents the probability that the nth sample is initially determined to be a real target echo signal, while represents the probability that it is determined to be a deception interference signal;
[0039] S44. Define the training objective of the deep learning recognition network as minimizing the cross-entropy loss function L(θ):
[0040]
[0041] where N represents the total number of training samples, y(n) is the true class label of the nth sample, y(n)=1 represents a real target echo signal, and y(n)=0 represents a deception interference signal;
[0042] S45. Use the gradient descent method to iteratively optimize the parameters θ of the deep learning recognition network D(θ) until the loss function L(θ) converges, so as to obtain the trained deep learning recognition network D(θ * );
[0043] S46. Use the trained deep learning recognition network D(θ * ) to preliminarily classify the input preprocessed radar echo signal data X processed (n) and output the preliminary classification result of the signal.
[0044] Optionally, the S5 specifically includes:
[0045] S51. Build an adaptive recognition optimization framework driven by reinforcement learning, integrate the trained deep learning recognition network D(θ * ) with an adaptive risk game model to form a joint optimization framework. The joint optimization framework comprehensively uses the preprocessed radar echo signal data X processed (n) for feature extraction, and combines the game optimization strategy p R and the risk return for dynamic adjustment to improve the adaptability to counter deception interference signals. The output of the joint optimization framework is described by the following formula:
[0046]
[0047] where D(θ* , X processed represents the spoofing interference features extracted by the deep learning recognition network, represents the comprehensive benefit of the risk game optimization strategy, σ(·) is the non-linear activation function, W a , W b is the weight matrix, and b is the bias term;
[0048] S52. Construct the state space of the reinforcement learning module, and define the system state s t as the comprehensive recognition environment of the radar system at the current time t:
[0049]
[0050] S53. Define the action space of the reinforcement learning. The action space is how the radar system dynamically adjusts the recognition strategy and the parameters of the deep learning model through reinforcement learning. The action a t is generated by the policy function π(s t ; φ):
[0051]
[0052] where, Δp R represents the adjustment amount of the current radar recognition strategy p R , including the response mode to the spoofing interference signal, and Δθ represents the optimization amount of the parameters θ of the deep learning recognition network. π(s t ; φ) represents the policy function of the reinforcement learning module;
[0053] S54. Construct a reward function to optimize the recognition strategy. The reward function r(s t , a t ) measures the improvement of the radar recognition ability and is calculated in combination with the risk benefit:
[0054]
[0055] where, ΔP acc represents the change in the radar recognition accuracy rate after taking the action a t , that is, the improvement degree of the correct recognition rate of the radar for the spoofing interference signal. ΔP fa represents the change in the false alarm rate, that is, the change in the probability that the radar system misjudges the spoofing interference as a real target. ΔP md represents the change in the missed alarm rate, that is, the change in the probability that the radar system misjudges a real target as a spoofing interference signal, represents the risk-adjusted benefit calculated after taking the action a t . α represents the minimum risk benefit threshold acceptable to the radar system. If Below α, the system policy will be automatically adjusted. κ is the risk adjustment factor, which controls the risk weight in extreme scenarios. w4, w5, and w6 are the weight coefficients of the recognition accuracy rate, false alarm rate, and missed alarm rate respectively;
[0056] S55. Policy optimization of the reinforcement learning module. The policy gradient method is used to optimize the reinforcement learning policy parameters φ, enabling the system to improve the intelligence of the radar recognition policy during the game process:
[0057]
[0058] Among them, η is the learning rate, which controls the parameter adjustment step size of the reinforcement learning module, represents the policy gradient;
[0059] S56. Synchronously optimize the radar recognition policy p R and the deep learning model parameters θ. Combining the output signal of the reinforcement learning, adopt the gradient update strategy:
[0060]
[0061] Among them, η R is the recognition policy update step size, which controls the optimization of the recognition policy p R in the adaptive risk game model. η θ is the deep learning network parameter update step size;
[0062] S57. State transition mechanism of the system. With the dynamic optimization of the radar recognition policy and the deep learning model, the system state is updated as follows:
[0063] s t+1 = T(s t , a t , r(s t , a t ));
[0064] Among them, T(·) is the system state transition function.
[0065] The beneficial effects of the present invention are:
[0066] (1) By introducing an adaptive risk game model, the present invention constructs an adversarial relationship between the radar system and deception jamming, and uses the Nash equilibrium to solve and optimize the radar recognition policy, thereby improving the anti-jamming ability of the radar system in a complex electromagnetic environment. The present invention can adjust the recognition policy of the radar system according to the dynamic changes of the combat environment, enabling it to maintain a high recognition accuracy rate when facing different types of deception jamming. In addition, through the optimized calculation of the risk-benefit function, the present invention can reduce the false alarm rate and missed alarm rate while ensuring the recognition accuracy.
[0067] (2) The present invention adopts a deep learning recognition network combined with a reinforcement learning optimization mechanism. The deep learning network automatically extracts the time-domain and frequency-domain features of radar echo signals, and the reinforcement learning module dynamically adjusts the recognition strategy and model parameters in different combat environments to achieve adaptive optimization. Traditional deep learning methods usually adopt a static training method. Once the training is completed, the model parameters are fixed, making it difficult to cope with the changes in deception interference strategies. By introducing the reinforcement learning module, the radar recognition strategy of the present invention can continuously learn through interaction with the environment and adjust the parameters of the deep learning recognition network according to the reward feedback, enabling the model to maintain high-efficiency recognition ability in complex electromagnetic environments.
[0068] (3) The present invention introduces a conditional value-at-risk calculation mechanism to model and optimize the decision-making risk of the radar system, enabling the radar to maintain better recognition performance under extreme interference conditions. Traditional deception interference recognition methods mainly rely on fixed risk control strategies, which are prone to serious false alarm or missed alarm problems under high deception interference intensity or unknown interference strategies. By calculating the maximum potential loss of the radar system in real time and using it as a constraint for strategy optimization, the radar system of the present invention can effectively avoid high-risk strategies, thereby reducing the uncertainty in combat decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0070] Figure 1 is a flowchart of a radar deception interference recognition method based on adversarial game optimization proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0072] Refer to Figure 1 , a radar deception interference recognition method based on adversarial game optimization, comprising the following steps:
[0073] S1. Obtain radar echo signal data, including real target echo signals and deception interference signals;
[0074] S2. Perform noise suppression, clutter suppression, and time-domain and frequency-domain feature extraction on the radar echo signal data to obtain preprocessed radar echo signal data;
[0075] S3. Construct an adaptive risk game model between the radar and deception interference based on the preprocessed radar echo signal data;
[0076] S4. Train a deep learning recognition network using the preprocessed radar echo signal data. The deep learning recognition network is used to extract the hidden features of the spoofing interference signal, and realize the preliminary classification of the spoofing interference signal and the real target echo signal;
[0077] S5. Introduce a reinforcement learning module, integrate the deep learning recognition network and the adaptive risk game model. The reinforcement learning module obtains a reward signal through interaction with the combat environment, and adaptively adjusts the radar recognition strategy and the parameters of the deep learning recognition network according to the reward signal;
[0078] S6. Apply the optimized radar recognition strategy to the processing of real-time radar echo signal data, identify and classify the spoofing interference signal in real time, and output the corresponding recognition results.
[0079] In this embodiment, S2 specifically includes:
[0080] S21. Sample the radar echo signal data x(t). The radar echo signal data x(t) consists of the real target echo signal s(t), the spoofing interference signal j(t), and the noise n(t);
[0081] S22. Perform noise suppression, and use an adaptive filtering method to reduce the Gaussian white noise interference to obtain the filtered radar echo signal data;
[0082] S23. Perform clutter suppression, enhance the real target signal component, reduce the ground clutter and sea clutter interference, and obtain the radar echo signal data after clutter suppression;
[0083] S24. Extract the time-domain and frequency-domain features, including short-time energy, instantaneous frequency, and time-frequency joint distribution, and obtain the preprocessed radar echo signal data X processed (n).
[0084] In this embodiment, S3 specifically includes:
[0085] S31. Construct an adaptive risk game model for the radar system and spoofing interference. Let the radar system be the game player R, and the spoofing interference be the opponent J. Among them, the radar system R distinguishes the real target echo signal and the spoofing interference signal through the optimal recognition strategy. The spoofing interference J uses the interference strategy to confuse the radar system and make it misidentify. Obtain the preprocessed radar echo signal data X processed (n) as the basis for the radar system to make a decision on the recognition strategy in the game. Combining the time-domain and frequency-domain features in the radar echo signal data, define the recognition strategy set of the radar system as S R , and define the interference strategy set of the spoofing interference as S J :
[0086]
[0087] Among them, represents the i-th recognition strategy of the radar system, represents the j-th interference strategy of the deception interference, and m and n are the number of strategies of the radar system and the deception interference respectively;
[0088] S32. Establish a revenue function based on the preprocessed radar echo signal data, and define the revenue function of the radar system and the revenue function of the deception interference The two satisfy the zero-sum game relationship:
[0089]
[0090] Among them, the revenue function of the radar system is driven by the feature data extracted from the preprocessed radar echo signal data and depends on the recognition accuracy rate P acc , the false alarm rate P fa and the miss alarm rate P md , the revenue function of the deception interference reflects its ability to successfully deceive the radar system:
[0091]
[0092] Among them, w1, w2, and w3 are the weight coefficients of the recognition accuracy rate, the false alarm rate, and the miss alarm rate respectively;
[0093] S33. Introduce an adaptive risk control mechanism, calculate the risk revenue of the radar system under different strategy combinations based on the radar echo signal data, use the conditional value at risk to calculate the maximum loss borne by the radar system, optimize the recognition strategy and reduce the false alarm and miss alarm of the risk situation above the risk threshold. The risk threshold of the radar system under the conditional value at risk is α, and the risk revenue is calculated as follows:
[0094]
[0095] Among them, represents the expected revenue of the radar system;
[0096] S34. Optimize the game strategy and solve the Nash equilibrium. Use the game optimization method to calculate the optimal mixed recognition strategy of the radar system and the optimal mixed interference strategy of the deception interference Solve the mixed strategy Nash equilibrium point that satisfies the following relationship:
[0097]
[0098] S35. The calculated optimal recognition strategy Applied to real-time radar signal processing, combined with radar echo signal data X processed (n) Real-time identify and classify deceptive interference signals, and output the final identification result. Through the risk adaptive adjustment mechanism, if the risk-benefit is lower than the set threshold α during the identification process, the radar system automatically adjusts the identification strategy.
[0099] In this embodiment, the recognition accuracy rate is the probability that the radar successfully distinguishes the real echo signal from the deceptive interference signal, the false alarm rate is the probability that the radar misidentifies the deceptive interference as a real target, and the missed alarm rate is the probability that the radar misidentifies the real target as a deceptive interference.
[0100] In this embodiment, S4 specifically includes:
[0101] S41. Use the preprocessed radar echo signal data X processed (n) as the input of the deep learning recognition network D(θ), where θ is the network parameter;
[0102] S42. Design the deep learning recognition network D(θ). The deep learning recognition network includes a feature extraction layer and a classifier layer. The feature extraction layer is used to extract the hidden features of the deceptive interference signal from the preprocessed radar echo signal data and output the feature vector f(n);
[0103] S43. Use the classifier layer to map the extracted feature vector f(n) to the class probability where The value range of is [0,1], indicating the probability that the nth sample is initially determined to be a real target echo signal, while Indicates the probability that it is determined to be a deceptive interference signal;
[0104] S44. Define the training objective of the deep learning recognition network as minimizing the cross-entropy loss function L(θ):
[0105]
[0106] where N represents the total number of training samples, y(n) is the true class label of the nth sample, y(n)=1 represents the real target echo signal, and y(n)=0 represents the deceptive interference signal;
[0107] S45. Use the gradient descent method to iteratively optimize the parameter θ of the deep learning recognition network D(θ) until the loss function L(θ) converges, so as to obtain the trained deep learning recognition network D(θ * );
[0108] S46. Use the trained deep learning recognition network D(θ *)Preprocess the input radar echo signal data X processed (n) for preliminary classification and output the preliminary classification result of the signal.
[0109] In this embodiment, S5 specifically includes:
[0110] S51. Construct an adaptive recognition optimization framework driven by reinforcement learning, integrate the trained deep learning recognition network D(θ * ) with the adaptive risk game model to form a joint optimization framework. The joint optimization framework comprehensively uses the preprocessed radar echo signal data X processed (n) for feature extraction, and combines the game optimization strategy p R and the risk and return for dynamic adjustment to improve the adaptability to counter deception interference signals. The output of the joint optimization framework is described by the following formula:
[0111]
[0112] Among them, D(θ * ,X processed (n)) represents the deception interference features extracted by the deep learning recognition network, represents the comprehensive return of the risk game optimization strategy, σ(·) is a non-linear activation function, W a ,W b is the weight matrix, and b is the bias term;
[0113] S52. Construct the state space of the reinforcement learning module and define the system state s t as the comprehensive recognition environment of the radar system at the current time t:
[0114]
[0115] S53. Define the action space of the reinforcement learning. The action space is how the radar system dynamically adjusts the recognition strategy and the parameters of the deep learning model through reinforcement learning. The action a t is generated by the policy function π(s t ; φ):
[0116]
[0117] Among them, Δp R represents the adjustment amount of the current radar recognition strategy p R , including the response mode to the deception interference signal, Δθ represents the optimization amount of the parameters θ of the deep learning recognition network, and π(s t ; φ) represents the policy function of the reinforcement learning module;
[0118] S54. Construct a reward function to optimize the recognition strategy. The reward function r(s t ,a t ) measures the improvement in the radar's recognition ability and is calculated in combination with risk and return:
[0119]
[0120] where ΔP acc represents the change in the radar's recognition accuracy after taking action a t , that is, the degree of improvement in the correct recognition rate of the radar for deceptive interference signals. ΔP fa represents the change in the false alarm rate, that is, the change in the probability that the radar system misjudges deceptive interference as a real target. ΔP md represents the change in the missed alarm rate, that is, the change in the probability that the radar system misjudges a real target as a deceptive interference signal. represents the risk-adjusted return calculated after taking action a t . α represents the minimum risk-return threshold acceptable to the radar system. If is lower than α, the system strategy will be automatically adjusted. κ is the risk adjustment factor, which controls the risk weight in extreme scenarios. w4, w5, and w6 are the weight coefficients of the recognition accuracy, false alarm rate, and missed alarm rate respectively;
[0121] S55. Policy optimization of the reinforcement learning module. The policy gradient method is used to optimize the reinforcement learning policy parameters φ, enabling the system to improve the intelligence of the radar recognition strategy during the game process:
[0122]
[0123] where η is the learning rate, which controls the parameter adjustment step size of the reinforcement learning module. represents the policy gradient;
[0124] S56. Synchronously optimize the radar recognition strategy p R and the deep learning model parameters θ. Combining the output signal of the reinforcement learning, a gradient update strategy is adopted:
[0125]
[0126] where η R is the recognition strategy update step size, which controls the optimization of the recognition strategy p R in the adaptive risk game model. η θ is the deep learning network parameter update step size;
[0127] S57. System state transition mechanism. As the radar recognition strategy and the deep learning model are dynamically optimized, the system state is updated as follows:
[0128] s t+1 = T(s t , a t , r(s t , a t ));
[0129] Among them, T(·) is the system state transition function.
[0130] Example 1:
[0131] At 02:30 in the early morning of June 15, 2024, at the naval radar monitoring station in a certain sea area of the East China Sea, the radar system began to perform routine airspace patrol and monitoring tasks. This radar station is equipped with a high-precision phased array radar, responsible for monitoring air targets within a range of 200 kilometers to prevent surprise attacks by enemy fighter jets or drones. At 02:45 in the early morning, the radar system detected an unidentified flying object in the sea area direction of 122.45° east longitude and 27.30° north latitude. This target was initially judged to be an enemy reconnaissance aircraft. The radar recorded its echo characteristics as follows:
[0132] Initial echo intensity: 47 dB; initial speed: 520 km / h; flight altitude: 8200 m; course angle: 215°; echo time delay: within the normal range (no abnormality detected).
[0133] However, at 02:50, the radar system detected that the echo time delay of this target began to drift abnormally. The target echo time delay increased by 10 microseconds, and the signal intensity showed a slight attenuation (dropped to 43 dB). Although this change is caused by target maneuvering under normal circumstances, the spectral feature analysis module of the radar system detected that the instantaneous frequency curve of this target showed a non-linear drift within the past 20 seconds and generated characteristics similar to Doppler repeater jamming at multiple time points.
[0134] At 02:55, the radar system used a deep learning recognition network to further analyze the echo data, extracted the short-time energy change characteristics, instantaneous frequency distribution, and time-frequency joint distribution of this target, and input the data into the radar deception jamming recognition model. After 0.8 seconds of calculation, the system determined that this target was suffering from delayed deception jamming because its echo characteristics had mutated twice within the past 30 seconds and the target's track showed abnormal drift.
[0135] At 02:58, the radar system detected a second target. The echo characteristics of this target were highly similar to those of the first target, but the position coordinates differed by 5.2 kilometers. At the same time, the system found that the echo signals of the two targets had a consistency as high as 97%, which means it was a repeater deception jamming - the enemy used a jammer to copy the radar echo and create a false target to interfere with the radar's judgment.
[0136] To further confirm, the radar system activates the adaptive risk game optimization module to calculate the optimal recognition strategy in the current game environment:
[0137] Current radar recognition strategy: SR3; Enemy deception jamming strategy: SJ2; Calculated optimal game strategy probability: pR* = 0.87.
[0138] At 03:00, the radar system finally confirms that the first target is a real fighter plane, while the second target is a false target created by the enemy jammer. The system then generates a deception jamming warning with the following content:
[0139] Deception jamming type: Repeater deception + Delay deception; Deception jamming start time: 02:50; Deception target: Radar station in the East China Sea area; Jammer azimuth: 122.87° E, 27.12° N; Matching degree of deception echo signal characteristics: 97%;
[0140] Target true probability assessment:
[0141] Target 1 (real fighter plane): 98.6%; Target 2 (deception signal): 1.4%.
[0142] The radar system sends this warning information to the naval air defense command center through an encrypted channel and enables the adaptive interference suppression mode at 03:02, automatically adjusting the radar waveform and signal processing parameters to reduce the impact of deception jamming. At the same time, the system uploads the latest game optimization parameters to the radar recognition strategy library to ensure faster recognition in case of similar interference in the future.
[0143] At 03:10, after the air defense command center confirms the target recognition result, it orders two fighter planes to take off for interception and commands the ground air defense system to enter the first-level alert state. At 03:20, the fighter planes arrive at the target area and confirm that the target information provided by the radar is accurate, and finally successfully drive away the enemy reconnaissance plane.
[0144] To further verify the effectiveness of the present invention, we conduct experiments in multiple real and simulated environments and compare the performance of the traditional method and the method of the present invention in terms of recognition accuracy, false alarm rate, and response time, as shown in Tables 1 - 4 below:
[0145] Table 1 Comparison of recognition accuracy between the traditional method and the method of the present invention
[0146]
[0147] Table 2 Adaptability of the traditional method and the method of the present invention in complex environments
[0148]
[0149] Table 3 Comparison of false alarm rates between the traditional method and the method of the present invention
[0150] Interference type False alarm rate of traditional method (%) False alarm rate of the method of the present invention (%) Delayed deception interference 12.4 4.1 Repeater deception interference 15.7 3.5 Intelligent interference 18.3 5.2
[0151] Table 4 Comparison of Recognition Time between Traditional Methods and the Method of the Present Invention
[0152]
[0153] In this embodiment, the recognition process of the radar system in a complex deception interference environment is successfully simulated, and it is shown how the method of the present invention combines deep learning, game optimization, and reinforcement learning to adaptively adjust the recognition strategy and improve the anti-interference ability of the radar system. Experimental data show that the method of the present invention is significantly superior to traditional methods in terms of recognition accuracy, false alarm rate, and recognition time. Especially in a high deception interference environment, it can still maintain a recognition accuracy of 94.1%, which is 20.5% higher than traditional methods, significantly enhancing the battlefield survival ability of the radar system.
[0154] The present invention improves the anti-interference ability of the radar system in a complex electromagnetic environment by introducing an adaptive risk game model, constructing an adversarial relationship between the radar system and deception interference, and using the Nash equilibrium to solve and optimize the radar recognition strategy. The present invention can adjust the recognition strategy of the radar system according to the dynamic changes of the combat environment, enabling it to maintain a high recognition accuracy when facing different types of deception interference. In addition, through the optimized calculation of the risk-reward function, the present invention can reduce the false alarm rate and missed alarm rate while ensuring recognition accuracy.
[0155] The present invention adopts a deep learning recognition network combined with a reinforcement learning optimization mechanism. The deep learning network automatically extracts the time-domain and frequency-domain features of radar echo signals, and the reinforcement learning module dynamically adjusts the recognition strategy and model parameters in different combat environments to achieve adaptive optimization. Traditional deep learning methods usually adopt a static training method. Once the training is completed, the model parameters are fixed, making it difficult to cope with changes in deception interference strategies. By introducing the reinforcement learning module, the radar recognition strategy of the present invention can continuously learn through interaction with the environment and adjust the parameters of the deep learning recognition network according to the reward feedback, enabling the model to maintain high recognition ability in a complex electromagnetic environment.
[0156] The present invention introduces a conditional value-at-risk calculation mechanism to model and optimize the decision-making risk of the radar system, enabling the radar to maintain better recognition performance under extreme interference conditions. Traditional deception interference recognition methods mainly rely on fixed risk control strategies, which are prone to serious false alarm or missed alarm problems under high deception interference intensity or unknown interference strategies. By calculating the maximum potential loss of the radar system in real time and using it as a constraint for strategy optimization, the present invention enables the radar system to effectively avoid high-risk strategies, thereby reducing the uncertainty in combat decision-making.
[0157] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered within the protection scope of the present invention.
Claims
1. A radar deception jamming recognition method based on adversarial game optimization, characterized in that It includes the following steps: S1. Obtain radar echo signal data, which includes real target echo signals and spoofing interference signals; S2. Perform noise suppression, clutter suppression, and time-domain and frequency-domain feature extraction on the radar echo signal data to obtain preprocessed radar echo signal data; S3. Construct an adaptive risk game model between the radar and spoofing interference based on the preprocessed radar echo signal data; S4. Use the preprocessed radar echo signal data to train a deep learning recognition network, which is used to extract the hidden features of spoofing interference signals and achieve a preliminary classification of spoofing interference signals and real target echo signals; S5. Introduce a reinforcement learning module to integrate the deep learning recognition network and the adaptive risk game model. The reinforcement learning module obtains a reward signal through interaction with the combat environment and adaptively adjusts the radar recognition strategy and the parameters of the deep learning recognition network according to the reward signal; S6. Apply the optimized radar recognition strategy to the processing of real-time radar echo signal data, identify and classify spoofing interference signals in real time, and output corresponding recognition results.
2. The radar deception jamming recognition method based on adversarial game optimization according to claim 1, wherein, The specific content of S2 includes: S21. Sample the radar echo signal data x(t), where the radar echo signal data x(t) consists of a real target echo signal s(t), a spoofing interference signal j(t), and noise n(t); S22. Perform noise suppression, adopt an adaptive filtering method to reduce Gaussian white noise interference, and obtain filtered radar echo signal data; S23. Perform clutter suppression, enhance the real target signal component, and reduce ground clutter and sea clutter interference to obtain radar echo signal data after clutter suppression; S24. Extract time-domain and frequency-domain features, including short-time energy, instantaneous frequency, and time-frequency joint distribution, and obtain the preprocessed radar echo signal data X processed (n).
3. The radar deception jamming recognition method based on adversarial game optimization according to claim 1, characterized in that The specific content of S3 includes: S31. Construct an adaptive risk game model for the radar system and deception jamming. Let the radar system be the game player R, and the deception jamming be the adversary J. Among them, the radar system R distinguishes the real target echo signal from the deception jamming signal through the optimal recognition strategy, and the deception jamming J confuses the radar system through the jamming strategy to make it misidentify, and obtains the preprocessed radar echo signal data X processed (n) is used as the basis for the radar system to decide the recognition strategy in the game. Combining the time-domain and frequency-domain features in the radar echo signal data, the recognition strategy set of the radar system is defined as S R , and the jamming strategy set of the deception jamming is defined as S J : Among them, represents the i-th recognition strategy of the radar system, represents the j-th interference strategy of the deception jamming, where m and n are the numbers of strategies of the radar system and the deception jamming respectively; S32. Establish a revenue function based on the preprocessed radar echo signal data and define the revenue function of the radar system and the revenue function of the deception jamming The two satisfy the zero-sum game relationship: Among them, the revenue function of the radar system is driven by the feature data extracted from the preprocessed radar echo signal data and depends on the recognition accuracy rate P acc , the false alarm rate P fa and the missed alarm rate P md . The revenue function of the spoofing jamming reflects its ability to successfully spoof the radar system: Among them, w1, w2, and w3 are the weight coefficients of the recognition accuracy rate, false alarm rate, and missed alarm rate respectively; S33. Introduce an adaptive risk control mechanism, calculate the risk benefit of the radar system under different strategy combinations based on the radar echo signal data, use conditional value at risk to calculate the maximum loss borne by the radar system, optimize the recognition strategy and reduce false alarms and missed alarms in risk situations above the risk threshold. The risk threshold of the radar system under conditional value at risk is α, and the risk benefit is calculated as follows: Among them, represents the expected revenue of the radar system; S34. Optimize the game strategy and solve the Nash equilibrium, and use the game optimization method to calculate the optimal mixed recognition strategy of the radar system and the optimal mixed jamming strategy of deception jamming Solve the mixed strategy Nash equilibrium point that satisfies the following relationship: S35. Apply the calculated optimal recognition strategy to real-time radar signal processing, and combine with the radar echo signal data X processed (n) to identify and classify deceptive interference signals in real time, and output the final recognition result. Through the risk adaptive adjustment mechanism, if the risk benefit is lower than the set threshold α during the recognition process, the radar system automatically adjusts the recognition strategy.
4. The radar deception jamming recognition method based on adversarial game optimization according to claim 3, characterized in that The recognition accuracy rate is the probability that the radar successfully distinguishes real echo signals from spoofing interference signals. The false alarm rate is the probability that the radar misidentifies spoofing interference as a real target. The missed alarm rate is the probability that the radar misidentifies a real target as spoofing interference.
5. The radar deception jamming recognition method based on adversarial game optimization according to claim 1, characterized in that, The specific content of S4 includes: S41. Using the preprocessed radar echo signal data X processed (n) as the input of the deep learning recognition network D(θ), where θ is the network parameter; S42. Design a deep learning recognition network D(θ), which includes a feature extraction layer and a classifier layer. The feature extraction layer is used to extract the hidden features of spoofing interference signals from the preprocessed radar echo signal data and output a feature vector f(n); S43. Use a classifier layer to map the extracted feature vector f(n) to class probabilities where ranges from [0, 1], representing the probability that the nth sample is preliminarily determined to be a real target echo signal, while represents the probability that it is determined to be a deception interference signal; S44. Define the training objective of the deep learning recognition network as minimizing the cross-entropy loss function L(θ): Among them, N represents the total number of training samples, y(n) is the true class label of the nth sample, y(n)=1 represents a real target echo signal, and y(n)=0 represents a spoofing interference signal; S45. The parameters θ of the deep learning recognition network D(θ) are iteratively optimized using the gradient descent method until the loss function L(θ) converges, so as to obtain the trained deep learning recognition network D(θ * ); S46. Using the trained deep learning recognition network D(θ * ) to perform preliminary classification on the preprocessed radar echo signal data X processed (n), and output the preliminary classification result of the signal.
6. The radar deception jamming recognition method based on adversarial game optimization according to claim 1, wherein The specific content of S5 includes: S51. Construct an adaptive recognition optimization framework driven by reinforcement learning, integrate the trained deep learning recognition network D(θ * ) with the adaptive risk game model to form a joint optimization framework. The joint optimization framework comprehensively utilizes the preprocessed radar echo signal data X processed (n) for feature extraction, and combines the game optimization strategy p R and the risk reward for dynamic adjustment to improve the adaptability to counter deception interference signals. The output of the joint optimization framework is described by the following formula: Among them, D(θ * , X processed (n)) represents the spoofing interference features extracted by the deep learning recognition network, represents the comprehensive benefit of the risk game optimization strategy, σ(·) is the non-linear activation function, W a , W b is the weight matrix, and b is the bias term; S52. Construct the state space of the reinforcement learning module and define the system state s t As the comprehensive recognition environment of the radar system at the current moment t: S53. Define the action space of reinforcement learning. The action space is how the radar system dynamically adjusts the recognition strategy and deep learning model parameters through reinforcement learning, and the action is a t Generated by the policy function π(s t ; φ): Among them, Δp R represents the adjustment amount for the current radar recognition strategy p R , including the response mode to the spoofing interference signal. Δθ represents the optimization amount for the deep learning recognition network parameter θ. π(s t ; φ) represents the policy function of the reinforcement learning module; S54. Construct a reward function to optimize the recognition strategy. The reward function r(s t , a t ) measures the improvement of the radar recognition ability and is calculated by combining risk and return: Among them, ΔP acc represents the change in the radar recognition accuracy after taking action a t , that is, the improvement degree of the correct recognition rate of the radar for the spoofing interference signal. ΔP fa represents the change in the false alarm rate, that is, the change in the probability that the radar system misjudges the spoofing interference as a real target. ΔP md represents the change in the miss alarm rate, that is, the change in the probability that the radar system misjudges a real target as a spoofing interference signal. represents the risk-adjusted return calculated after taking action a t . α represents the minimum risk return threshold acceptable to the radar system. If is lower than α, the system strategy will be automatically adjusted. κ is the risk adjustment factor to control the risk weight in extreme scenarios. w4, w5, and w6 are the weight coefficients of the recognition accuracy, false alarm rate, and miss alarm rate respectively; S55. Policy optimization of the reinforcement learning module, using the policy gradient method to optimize the reinforcement learning policy parameters φ, so that the system can improve the intelligence of the radar recognition strategy during the game process: where η is the learning rate, which controls the parameter adjustment step size of the reinforcement learning module, represents the policy gradient; S56. Synchronously optimize the radar recognition strategy p R Combine with the deep learning model parameter θ and the output signal of reinforcement learning, and adopt a gradient update strategy: Among them, η R is the recognition strategy update step size, which controls the optimization of the recognition strategy p R in the adaptive risk game model, and η θ is the deep learning network parameter update step size; S57. State transition mechanism of the system. As the radar recognition strategy and the deep learning model are dynamically optimized, the system state is updated as follows: s t+1 = T(s t , a t , r(s t , a t )); where T(·) is the system state transition function.
Citation Information
Cited By
Radar seeker anti-interference method based on detection and detection integration
CN120630120A
A radar seeker anti-jamming method based on integrated detection and tracking
CN120630120B
Radar adaptive anti-interference decision method based on deep learning
CN121541150A