Q-learning based multi-function radar jamming decision method

CN118033558BActive Publication Date: 2026-09-29XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410344946.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-25
Publication Date
2026-09-29
Estimated Expiration
2044-03-25

AI Technical Summary

Benefits of technology

[0019](1)本发明根据当前时刻雷达工作模式的识别结果及其对应的对雷达进行干扰的评估结果对动作价值函数进行更新,充分考虑了评估结果对准确率的影响,避免了现有技术仅使用雷达的威胁等级导致信息缺失的缺陷,有效提高了决策的准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118033558B_ABST
    Figure CN118033558B_ABST
Patent Text Reader

Abstract

The application provides a multifunctional radar jamming decision method based on Q learning, and the implementation steps are as follows: obtaining a training data set and to-be-identified data at a current moment; identifying the working mode of the radar; selecting a jamming mode and evaluating the jamming effect; updating the action value function; and obtaining a jamming strategy.The application updates the action value function according to the identification result of the working mode of the radar at the current moment and the evaluation result of the radar jamming corresponding to the identification result, fully considers the influence of the evaluation result on the accuracy, avoids the defects of the prior art that only uses the threat level of the radar to cause information loss, effectively improves the accuracy of the decision, and uses a dynamic exploration method to select the jamming mode; with the continuous interaction and confrontation with the radar, the exploration probability decreases from large to small, avoids the influence of the fixed exploration probability of the prior art on the convergence speed of the algorithm, and effectively improves the efficiency of the decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of radar countermeasures technology and relates to a radar jamming decision-making method, specifically a multi-functional radar jamming decision-making method based on Q-learning. Background Technology

[0002] Multifunctional radar is an advanced type of radar with multiple operating modes such as search, track, identification, and guidance. Moreover, it can flexibly adjust radar parameters and operating modes according to changes in the battlefield threat situation, which brings great challenges to radar jamming decision-making.

[0003] Radar jamming decision-making refers to the process by which a jammer, based on reconnaissance radar information, selects the appropriate jamming pattern by observing changes in the enemy radar's operating mode. Traditional radar jamming decision-making methods rely on a large amount of prior knowledge or matching based on expert systems. These methods are unable to make rapid decisions against multi-functional radars that flexibly adjust their operating modes, and the jamming resource database cannot be updated in real time, making it difficult to adapt to the rapidly changing demands of the battlefield.

[0004] The Q-learning-based multi-functional radar jamming decision-making method refers to the process by which a jammer dynamically interacts with the radar, acquires feedback information, and continuously adjusts its decisions. It mainly includes the following steps: Cognitive Recognition: The jammer needs to acquire information from the battlefield environment and recognize the radar's operating mode. Jamming Decision: This refers to the jammer judging and deciding on the appropriate jamming pattern based on the radar's operating mode information. Closed-Loop Feedback: During the decision-making process, the jammer adjusts and updates its decisions based on the jamming effect evaluation as feedback information.

[0005] For example, patent application CN114415126A, entitled "A Radar Suppression Jamming Decision-Making Method Based on Reinforcement Learning," discloses a radar suppression jamming decision-making method based on reinforcement learning. This invention first constructs a radar state set based on the radar's operating mode and waveform parameters, and then constructs a jamming action set based on the jamming pattern and jamming parameters. Next, it calculates the gain value based on the radar's threat level. Finally, it learns jamming decisions through Q-learning and implements the optimal jamming action against the radar using a jammer. This invention has the technical characteristics of high efficiency and high accuracy, but because it only calculates the gain value based on the radar's threat level, the accuracy of the decision-making is still relatively poor; and because it selects the jamming pattern using a fixed exploration probability, the convergence efficiency of the algorithm is still relatively low. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the existing technology and propose a multi-functional radar jamming decision-making method based on Q-learning to solve the technical problems of low decision-making accuracy and efficiency in the existing technology.

[0007] To achieve the above objectives, the technical solution adopted by the present invention includes the following steps:

[0008] (1) Obtain the training dataset and the data to be identified at the current time:

[0009] The feature data of M radar signals contained in each of the L working modes of the multi-function radar, as well as the feature data to be identified at time t, are obtained and normalized respectively. Then, each working mode is used as the label of its corresponding normalized feature data. The normalized H = L × M feature data and their labels are combined to form a training sample set, and t = 1, where L ≥ 6, M ≥ 100, t ∈ [1, T], T is the total number of time points, T ≥ 1000.

[0010] (2) Identify the radar's operating mode:

[0011] Based on the K-nearest neighbor algorithm, and using the training sample set and the normalized feature data to be identified at time t, the radar's operating mode at the current time is identified, and the identification result S of the radar's operating mode at the current time is obtained. t ;

[0012] (3) Select the interference pattern and evaluate the interference effect:

[0013] The interference pattern J is selected using a dynamic exploration method. t And using the analytic hierarchy process and the Topsis algorithm, through interference pattern J t G signal parameters obtained at time t by jamming the radar The interference effect was evaluated, and the evaluation result R was obtained. t ;

[0014] (4) Update the action value function:

[0015] The Q-learning algorithm is employed, and the results of the radar operating mode identification S' after interference are used. t And evaluation results R t For the action value function Q(S) t J t The updated action value function Q'(S) is then used to update the action value. t J t ) Calculate the interference strategy π(J) at time t. t |S t );

[0016] (5) Obtaining interference strategies:

[0017] Determine whether t = T or Q'(S) t J t )-Q(S t J tIf )≤e holds true, then the final interference strategy π(J|S) is obtained; otherwise, let t=t+1, Q(S) t J t )=Q'(S t J t ), and perform step (2), where e≤0.0001.

[0018] Compared with the prior art, the present invention has the following advantages:

[0019] (1) The present invention updates the action value function based on the identification result of the radar working mode at the current moment and the corresponding evaluation result of the radar interference. It fully considers the impact of the evaluation result on the accuracy, avoids the defect of the prior art that only uses the radar threat level and causes information loss, and effectively improves the accuracy of decision-making.

[0020] (2) The present invention adopts a dynamic exploration method to select the interference pattern. As there is continuous interaction and confrontation with the radar, the exploration probability decreases from large to small. This avoids the impact of the fixed exploration probability on the convergence speed of the algorithm in the existing technology, and effectively improves the efficiency of decision-making.

[0021] (3) The present invention identifies the radar working mode based on the K-nearest neighbor algorithm, which avoids the disadvantage of directly obtaining the working mode from the radar in the prior art and is more practical. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating the implementation of the present invention.

[0023] Figure 2 This is a comparison chart of simulation results of the decision-making efficiency of the present invention and existing technologies. Detailed Implementation

[0024] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0025] Reference Figure 1 The present invention includes the following steps:

[0026] Step 1) Obtain the training dataset and the data to be identified at the current time:

[0027] To obtain the feature data of M radar signals contained in each of the L operating modes of the multi-functional radar, and the feature data to be identified at time t. To eliminate the influence of dimensions and orders of magnitude, all data need to be normalized. Then, each operating mode is used as the label of its corresponding normalized feature data, and the normalized H = L × M feature data and their labels are combined to form a training sample set, let t = 1, where S ≥ 6, M ≥ 100, t ∈ [1, T], T is the total number of time points, T ≥ 1000;

[0028] In this embodiment, L = 6, M = 100, and T = 4000.

[0029] In this embodiment, the radar has L operating modes, including: guidance mode, imaging mode, non-cooperative target identification mode, ranging mode, surveillance mode, and stereo search mode, with the threat level decreasing progressively.

[0030] The characteristic data of the radar signal includes N features: smoothness C, dispersion D, beam dwell number Ω, pulse width PW, duty cycle DC, and intra-pulse modulation type MT, where N=6.

[0031] The normalization formula for the h feature data included in each of the S operating modes of the multi-function radar is as follows:

[0032]

[0033] Among them, O h O' represents the h-th feature data. h O h The normalized result.

[0034] Step 2) Identify the radar's operating mode:

[0035] K-Nearest Neighbors (KNN) is an online technique that is theoretically simple and easy to implement. It classifies data by calculating the difference between the feature data to be identified and the training sample set. Generally, the distance between the data in each dimension is used to describe the difference between the feature data to be identified and the training sample set. Common methods include Euclidean distance, Mahalanobis distance, and Manhattan distance. In this embodiment, Euclidean distance is used to describe the difference.

[0036] Based on the K-nearest neighbor algorithm, the radar's operating mode at the current time is identified using the training sample set and the normalized feature data to be identified at time t, resulting in the identification result S of the radar's operating mode at the current time. t The implementation steps are as follows:

[0037] Calculate the Euclidean distance d(O) between each feature data in the training sample set and the feature data to be identified x. h The label that appears most frequently in the feature data corresponding to the smallest K Euclidean distances is taken as the recognition result S of the working mode of the data to be identified. t The formula for calculating d(Oh,x) is:

[0038]

[0039] Among them, O hn Let x be the nth dimension feature of the h-th feature data in the training sample set. n Let K be the nth dimension feature of x, and in this example, K = 9.

[0040] Step 3) Select the interference pattern and evaluate the interference effect:

[0041] (3a) Select the interference pattern:

[0042] With probability ε, one interference pattern J is randomly selected from the interference pattern set J = {J1, J2, J3, J4, J5, J6}. t Alternatively, the interference pattern J can be selected with a probability of 1-ε. t , Where J1, J2, J3, J4, J5, and J6 represent high-power signal suppression interference, smart noise interference, cross-eye interference, distance spoofing interference, noise modulation interference, and frequency shifting interference, respectively, and ε is the dynamic exploration probability, calculated using the following formula:

[0043]

[0044] Where f is a constant, f∈(0,0.1), and in this example f=0.05.

[0045] There are two probabilities when selecting interference patterns: ε represents the exploration probability, which randomly selects an interference pattern to ensure that the algorithm can explore and learn multiple times; 1-ε represents the utilization probability, which selects the radar operating mode at the current moment to optimize the action value function Q. π (S t J t The highest J t As a perturbation pattern for the next step, it ensures that the algorithm can utilize existing learning experience. Balancing the exploration probability and the exploitation probability allows the algorithm to converge quickly and complete the learning of decision-making.

[0046] The dynamic exploration probability ε is higher when t is small, allowing for the experimentation of different interference patterns to improve the action value function Q. π (S t J t In more different J t The following updates are made; as the exploration probability of t decreases, the exploration probability gradually decreases while the utilization probability gradually increases. At this point, the selected interference pattern has a higher probability of selecting the action value function Q. π (S t J t The highest J t It can help the action value function Q π (S t J t At the highest J t The process is updated multiple times to accelerate the convergence speed.

[0047] (3b) Obtain the interference pattern J tG signal parameters obtained at time t by jamming the radar

[0048]

[0049] in, It contains G parameters, namely: T f Indicates beam dwell time, Pri represents pulse repetition interval, N f Indicates the number of pulse groups, Ω f Indicates the frequency agility range, ν f Indicates frequency agility, Ω pri Indicates the agility range of the pulse repetition period, ν pri τ represents the pulse repetition period agility speed, τ represents the pulse width, and pw represents the bandwidth;

[0050] (3c) Calculate the weight vector of the parameters using the analytic hierarchy process:

[0051] The acquired G signal parameters can describe the behavior changes of radar signals after being interfered with. However, since the differences in the importance of each parameter can affect the evaluation results, weighting is required. Calculating the weights using the analytic hierarchy process (AHP) can transform subjective qualitative indicators into relative quantitative indicators for comparison and evaluation, ensuring the rationality and reliability of the weights.

[0052] Using the Analytic Hierarchy Process (AHP) to analyze signal parameters Each parameter in the algorithm is scored, and a judgment matrix A with a dimension of G×G is constructed using the nine score results V. Then, the weight ω of each parameter is calculated using A. g The weight vector W is formed, where:

[0053]

[0054]

[0055] W=(ω1,ω2,…,ω g ,…,ω G ) T

[0056] Among them, a ij V represents the i-th rating result in the judgment matrix A. i With the j-th rating result V j The ratio of , i∈[1,9], j∈[1,9], i≠j, (·) T This represents the transpose of a matrix. In this example, V = {1,3,2,4,4,4,4,1,1}.

[0057] (3d) The evaluation results were calculated using the Topsis algorithm:

[0058] The Topsis algorithm is a commonly used method for multi-attribute evaluation analysis. It evaluates solutions by calculating the distance between each solution and the positive and negative ideal solutions. The weights can be adjusted according to specific problems and needs, providing an objective standard for the comprehensive evaluation of various solutions. The steps for evaluating the interference effect using the Topsis algorithm are as follows:

[0059]

[0060]

[0061]

[0062]

[0063]

[0064]

[0065] Among them, Z tg express The weighted result of the g-th parameter, Let Z represent the g-th parameter Z at all times before time t. g The maximum and minimum values, Let Z represent the positive ideal solution composed of the weighted sum of all parameters and the maximum and minimum values ​​of all parameters before time t, respectively. + Negative ideal solution Z - The distance between them.

[0066] Interference assessment results are obtained through weighted signal parameters. The degree of proximity to the positive and negative ideal solutions is used as a quantitative evaluation value. The closer the solution is to the positive ideal solution and the farther it is from the negative ideal solution, the better the interference evaluation result.

[0067] Step 4) Update the action value function:

[0068] Q-learning is a reinforcement learning method that aims to explore how an agent can take optimal actions in complex and uncertain environments to maximize its rewards. Applying Q-learning to the field of radar countermeasures, the radar's operating mode corresponds to environmental information, the jamming pattern corresponds to the agent's actions, and the jamming effect corresponds to the environment's reward value. In Q-learning, an action-state value function Q(S,J) needs to be maintained to represent the value of different jamming patterns under the current radar operating mode. The higher the value, the more significant the use of the jamming pattern. Here, S represents the current radar operating mode, and J represents different jamming patterns.

[0069] The Q-learning algorithm is employed, and the results of the radar operating mode identification S' after interference are used. t And evaluation results R t For the action value function Q(S) t J t Update the function to obtain the updated action value function Q'(S). t J t ), and interference strategies The updated formula is:

[0070]

[0071] In this example, the learning rate α = 0.1 and the discount factor γ = 0.9.

[0072] When updating the action value function, R t The interference assessment result calculated in step 3) is used with R... t Updating can make full use of the effects of different interference patterns as feedback information, which helps to improve the accuracy of decision-making.

[0073] Step 5) Obtain the interference strategy:

[0074] Determine whether t = T or Q'(S) t J t )-Q(S t J t If ) ≤ e holds true, then the final interference strategy is obtained. Otherwise, let t = t + 1, Q(S t J t )=Q'(S t J t ), and perform step 2).

[0075] As t continuously increases, the action value function Q(S) t J t ) is continuously updated and gradually converges, when Q'(S t J t )-Q(S t J t If 0.0001 ≤ e, the system is considered converged. At this point, the jamming decision system has learned the optimal jamming pattern for each radar operating mode. In this example, e = 0.0001.

[0076] The technical effects of the present invention will be further explained below with reference to simulation experiments.

[0077] 1. Simulation conditions and content:

[0078] Software environment: Intel(R) Core(TM) i5-10300H CPU@2.50GHz, PyCharm emulation software under Windows 10 Home Chinese Edition 64-bit operating system.

[0079] The decision efficiency of this invention is compared with that of an existing radar suppression jamming decision-making method based on reinforcement learning through simulation. The results are as follows: Figure 2 As shown.

[0080] 2. Simulation Result Analysis:

[0081] Reference Figure 2 The top curve represents the sum of action value functions using the dynamic exploration probability of this invention as a function of time. The four curves below represent the sum of action value functions using the existing technology as a function of time when the exploration probabilities are 0, 0.3, 0.6, and 0.9, respectively. As can be seen from the figure, as t gradually increases, the sum of action value functions gradually increases and converges. When t = 1500, the dynamic exploration probability Q-learning algorithm achieves convergence, with faster convergence speed and higher efficiency.

Claims

1. A multi-functional radar jamming decision-making method based on Q-learning, characterized in that, Includes the following steps: (1) Obtain the training dataset and the data to be identified at the current time: Acquiring multi-functional radar Each of the various working modes includes Characteristic data of radar signals, and The feature data to be identified at each time step is normalized, and then each working mode is used as the label for its corresponding normalized feature data. The normalized data is then... The training sample set consists of feature data and their labels, and let... ,in, , , , The total number of moments. ; (2) Identify the radar's operating mode: based on Nearest neighbor algorithm, and through training sample set and The normalized feature data to be identified at each moment is used to identify the radar's operating mode at the current moment, thus obtaining the identification result of the radar's operating mode at the current moment. ; (3) Select the interference pattern and evaluate the interference effect: Selecting interference patterns using a dynamic exploration method It employs the analytic hierarchy process (AHP) and the Topsis algorithm to detect interference patterns. Acquired by jamming radar Moment individual signal parameters The interference effect was evaluated, and the evaluation results were obtained. Selecting the interference style The implementation steps are as follows: by To dynamically explore probabilities in the interference pattern set Randomly select an interference pattern , or Choose the interference pattern for probability ;in, , , , , , These represent high-power signal suppression interference, clever noise interference, crosseye interference, distance spoofing interference, noise modulation interference, and frequency shifting interference, respectively. The calculation formula is: ; in, It is a constant. , For An exponential function with base 0. This represents the parameter value that maximizes the function value; (4) Update the action value function: The Q-learning algorithm is employed, and the results are used to identify the radar's operating mode after interference. and evaluation results Action value function Update the function, and then use the updated action value function. calculate Interference strategies at different times ; (5) Obtaining interference strategies: judge or If the condition is met, then the final interference strategy is obtained. Otherwise, let , and perform step (2), where .

2. The method according to claim 1, characterized in that, The radar signal feature data and the feature data to be identified mentioned in step (1) both include One characteristic: smoothness Dispersion Beam dwell number Pulse width Duty cycle Intrapulse modulation type , .

3. The method according to claim 2, characterized in that, The normalized result described in step (1) The feature data, of which the first Feature data The normalization formula is: ; in, express The normalized result.

4. The method according to claim 3, characterized in that, The steps for identifying the radar's current operating mode as described in step (2) are as follows: Calculate the first training sample set Feature data With the feature data to be identified Euclidean distance and the smallest first The label that appears most frequently in the feature data corresponding to each Euclidean distance is used as the recognition result of the working mode of the data to be identified. ,in: ; in, , The first Feature data, feature data to be identified The Dimensional features.

5. The method according to claim 1, characterized in that, The interference pattern described in step (3) After jamming the radar Moment individual signal parameters The steps to evaluate the interference effect are as follows: (3a) Obtain the interference pattern After jamming the radar Moment individual signal parameters : ; in, They are: Indicates beam dwell time. Indicates the pulse repetition interval. Indicates the number of pulse groups. Indicates the frequency agility range. Indicates frequency agility speed. Indicates the agile range of the pulse repetition period. Indicates the pulse repetition period agility speed. Indicates pulse width. Indicates bandwidth; (3b) Using the Analytic Hierarchy Process (AHP) to analyze signal parameters Each parameter in the algorithm is scored, and nine score results are obtained from the scores. Construct the judgment dimension as Judgment matrix Then through Calculate the weight of each parameter , forming a weight vector ,in: ; ; ; in, Represents the judgment matrix The Middle Each rating result With the Each rating result The ratio, , , , Represents the transpose of a matrix; (3c) The Topsis algorithm is adopted, and the weight vector is used. Calculate the evaluation value of the interference effect : ; ; ; ; ; ; in, express The Middle The weighted result of each parameter , They represent All times before time 1 Parameters The maximum and minimum values, , Represent the weighted result of all parameters and The positive ideal solution is composed of the maximum and minimum values ​​of all parameters before time step [time]. Negative ideal solution The distance between them.

6. The method according to claim 1, characterized in that, The identification results of the radar operating mode after interference as described in step (4) and evaluation results Action value function The update is performed using the following formula: ; in, For learning rate, This is the discount factor.

7. The method according to claim 1, characterized in that, The updated action value function described in step (4) calculate Interference strategies at different times The calculation formula is: 。

Citation Information

Patent Citations

  • Radar suppressing jamming decision-making method based on reinforcement learning

    CN114415126A

  • Multifunctional radar cognitive interference decision-making method based on threat assessment

    CN116577739A

  • Method of improving a radar system, module for improving a radar system and an improved radar system

    US20230384415A1