Active phased array radar cognitive jamming signal generation method, medium and equipment

By constructing a radar electronic countermeasures simulation model and a DDQN network, a signal pattern state transition table is generated, which solves the shortcomings of traditional radar jamming decision-making methods under complex signal patterns, achieves more efficient jamming decision-making, and improves the performance and robustness of the radar electronic countermeasures system.

CN119758263BActive Publication Date: 2026-01-27CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411871928.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2026-01-27
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

Traditional radar jamming decision-making methods cannot cope with complex operating modes and flexible and ever-changing signal patterns. Jamming decisions based on reinforcement learning suffer from Q-value overestimation, lack adaptability to actual radar systems, and are difficult to implement effectively in airborne active phased array fire control radars.

Method used

A radar electronic countermeasures simulation model was constructed. A signal pattern state transition table was generated using the TOPSIS sorting algorithm and the minimum interference-to-signal ratio data table. The jamming pattern was selected by combining the DDQN network. The signal pattern transition table was constructed through a radar signal processing simulation system and actual radar experiments to form a countermeasures memory. The network parameters were updated by the DDQN algorithm, which solved the shortcomings of traditional methods.

Benefits of technology

It improves the accuracy and applicability of radar jamming decision-making, eliminates the problem of Q-value overestimation, enhances the accuracy and reliability of jamming decision-making, and meets the jamming decision-making requirements of complex airborne fire control radars.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119758263B_ABST
    Figure CN119758263B_ABST
Patent Text Reader

Abstract

The application discloses an active phased array system radar cognitive jamming signal generation method, medium and equipment, relates to the field of radar electronic countermeasure technology, and mainly comprises the following steps: constructing a radar electronic countermeasure simulation model, obtaining a signal pattern state conversion table by using a TOPSIS sorting algorithm and a minimum jamming signal ratio data table; obtaining an interference pattern selection model by using the signal pattern state conversion table and a DDQN network according to the radar electronic countermeasure simulation model; and performing optimal interference pattern decision by using the interference pattern selection model. The active phased array system radar cognitive jamming signal generation method, medium and equipment can improve the precision and applicability of radar jamming decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of radar electronic countermeasures technology, and more specifically, to a method, medium, and device for generating cognitive jamming signals for active phased array radar. Background Technology

[0002] In the field of radar electronic countermeasures, airborne active phased array (APA) fire control radars are widely used, such as in military fighter jets and missile defense systems. Currently, airborne APA fire control radars include typical signal patterns such as STT, TAS, TWS, RWS, and VS, while cognitive jammers include typical DRFM jamming signals such as comb spectrum jamming, sample jamming, slice jamming, Doppler noise jamming, random false target jamming, and delay-superimposed smart noise jamming. Traditional jamming decision-making methods may not be able to handle complex operating modes and flexible signal patterns. Current reinforcement learning-based jamming decision-making methods lack a reasonable basis for radar state transition relationships, relying on prior knowledge to construct transition relationship tables, lacking adaptation to actual radar systems. Furthermore, current DQN algorithms suffer from Q-value overestimation during the decision-making process, which affects the decision-making effect. Faced with the complex and varied signal patterns of airborne APA fire control radars, traditional experience-based jamming decisions are ineffective. In addition, with the continuous development of radar technology and the expansion of its application areas, the demand for more intelligent and adaptive jamming decision-making methods is increasing. Summary of the Invention

[0003] The purpose of this invention is to provide a method for generating cognitive jamming signals for active phased array radar, a storage medium, a computer device, and a computer program product, which can improve the accuracy and applicability of radar jamming decision-making.

[0004] This invention provides a method for generating cognitive jamming signals for active phased array radar, comprising the following steps: S1: Constructing a radar electronic countermeasures simulation model, and obtaining a signal pattern state transition table using the TOPSIS sorting algorithm and the minimum interference-to-signal ratio data table; S2: Obtaining a jamming pattern selection model based on the radar electronic countermeasures simulation model, using the signal pattern state transition table and the DDQN network; S3: Making an optimal jamming pattern decision using the jamming pattern selection model.

[0005] Further, step S1 specifically includes: S11: Constructing a radar electronic countermeasures simulation model, which includes an airborne active phased array radar model and a cognitive jammer model; S12: Obtaining error measurement results based on the radar electronic countermeasures simulation model; S13: Sorting the error measurement results using the TOPSIS sorting algorithm to obtain sorting results; S14: Correcting the sorting results using the minimum interference-to-signal ratio data table to obtain a corrected TOPSIS sorting table; S15: Obtaining a signal pattern state transition table based on the corrected TOPSIS sorting table using transformation rules.

[0006] Furthermore, the aforementioned airborne active phased array radar model includes the full angle of entry search (AAS) mode, high angle of entry search (HAS) mode, hybrid angle of entry search (CAM) mode, and single target tracking (STT) mode. The airborne active phased array radar model also includes a radar signal pattern set, which includes single target tracking (STT) radar signal patterns, search and track (TAS) radar signal patterns, search and track simultaneously (TWS) radar signal patterns, medium repetition rate (MRPR) search and ranging simultaneously (RWS-1) radar signal patterns, high repetition rate (MRPR) search and ranging simultaneously (RWS-2) radar signal patterns, and velocity search (VS) radar signal patterns. The cognitive jammer model includes a jamming signal pattern set, which includes comb spectrum jamming signal patterns, sample jamming signal patterns, Doppler noise jamming signal patterns, random false target jamming signal patterns, delay superposition agile noise jamming signal patterns, and slice jamming signal patterns.

[0007] Furthermore, the above error measurement results include ranging error, speed measurement error, false alarm probability, and detection probability.

[0008] Furthermore, step S13 specifically includes:

[0009] S131: Based on the error measurement results, the decision matrix is ​​obtained, as shown in the formula:

[0010]

[0011] Where K is the decision matrix, k nm Let be the parameter values ​​corresponding to the nth evaluation index in the mth reconnaissance result;

[0012] S132: Normalize the decision matrix to obtain the normalized matrix, as shown in the formula:

[0013]

[0014] Where, k ij k′ is the element in the i-th row and j-th column of the decision matrix. ij The element in the i-th row and j-th column of the normalized matrix;

[0015] S133: Weight the normalized matrix to obtain the weighted matrix, as shown in the formula:

[0016] z ij =w j k ij ′(i=1,2,L,m; j=1,2,L,n),

[0017] W = [w1, w2, L, w n ] T ,

[0018] Among them, z ij w is the element in the i-th row and j-th column of the weighted matrix; j Let J be the j-th element of the weight vector, and W be the weight vector.

[0019] S134: Based on the weighted matrix, the optimal reconnaissance result and the worst reconnaissance result are obtained, as shown in the formula:

[0020]

[0021] Among them, I + For optimal reconnaissance results, For the j-th element of the optimal reconnaissance result, I - As the worst reconnaissance result, Let j be the j-th element of the worst reconnaissance result, and max{·} represents taking the maximum value;

[0022] S135: Calculate the distance between each reconnaissance result and the best and worst reconnaissance results respectively, to obtain the first distance and the second distance, as shown in the formula:

[0023]

[0024] in, The first distance, This is the second distance;

[0025] S136: Based on the first and second distances, the quality of each reconnaissance result is determined, as shown in the formula:

[0026]

[0027] Among them, A i The quality of the i-th reconnaissance result;

[0028] S137: The ranking results are obtained based on the quality of each reconnaissance result.

[0029] Further, step S2 specifically includes: S21: Based on the radar electronic countermeasures simulation model, using the signal pattern state transition table and DDQN network, obtain adversarial samples; S22: Construct and update the memory bank based on the adversarial samples; S23: Use the memory bank as training samples to train the DDQN network and obtain the jamming pattern selection model.

[0030] Further, step S21 specifically includes: S211: Based on the radar electronic countermeasures simulation model, obtain a state set and an action set; the state set includes states, which represent the signal pattern of the enemy's airborne active phased array fire control radar at a certain moment; the action set includes actions, which represent the jamming signal pattern selected by our jammer at that moment; S212: Based on the state set and action set, use the signal pattern state transition table and DDQN network to obtain the Q value corresponding to each jamming signal pattern; S213: Based on the Q value corresponding to each jamming signal pattern, use the ε-greed strategy to obtain the first jamming signal pattern; S214: Based on the first jamming signal pattern, use the radar electronic countermeasures simulation model and the signal pattern state transition table to obtain the change of the enemy radar signal pattern, and obtain the first reward value based on the change of the enemy radar signal pattern; S215: Based on the first reward value, use the DDQN network to update the Q value, and use the signal pattern state transition table to obtain the next state; S216: Combine the state, the first jamming signal pattern, the first reward value, and the next state into a quadruple sample to obtain the adversarial sample.

[0031] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method for generating cognitive jamming signals for active phased array radar.

[0032] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the above-described method for generating cognitive jamming signals for active phased array radar.

[0033] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method for generating cognitive jamming signals for active phased array radar.

[0034] The active phased array radar cognitive jamming signal generation method, storage medium, computer equipment, and computer program product provided by this invention have the following beneficial effects:

[0035] This invention combines simulation data obtained from a radar electronic countermeasures simulation model with small-sample measured data to form a sufficiently effective signal pattern conversion table. The optimal interference decision scheme is obtained through the DDQN algorithm. The signal pattern conversion table is obtained using a radar signal processing simulation system and actual radar interference experiments. During the implementation of the signal pattern conversion table, the simulation system reads radar signal parameters from the simulated radar data, simulates using a signal-level simulation model, and adds different types of interference. Based on the results obtained from the PD / CFAR processing of the system before and after interference, it calculates indicators such as ranging error, velocity error, false alarm probability, and detection probability. The interference effect is then ranked according to the TOPSIS ranking algorithm. Specifically, the simulated radar system interference experiment is used to obtain the TOPSIS ranking results of target measurement errors, and the actual interference experiment is used to obtain the minimum interference-to-signal ratio for different interference patterns under different radar states. The minimum interference-to-signal ratio (MSR) data table is used to correct the topsis ranking table, and the signal pattern state transition is obtained based on the corrected topsis ranking table. Then, an adversarial memory is generated based on the signal pattern transition table, and finally, the optimal interference pattern is obtained through DDQN. The implementation of DDQN interference decision-making mainly relies on two neural networks to fit the Q-value: MainNet is used to estimate the Q-value, while TargetNet is used to generate the Q-target value. The network parameters are updated using the gradient descent algorithm based on the loss function formula. To improve the interference training effect, this invention employs an adversarial memory, which stores radar signal samples before and after interference, interference signal samples, and their corresponding interference effect values. The interference effect value is determined by comparing the changes in the signal before and after interference. During interference training, samples are randomly selected from the memory to avoid significant fluctuations in network fitting caused by the correlation between consecutive samples.

[0036] This invention addresses the challenge that traditional empirical jamming decisions are ineffective for airborne active phased array fire control radars under complex operating modes and variable signal patterns. Furthermore, it solves the problem of constructing the signal pattern conversion matrix for phased array radars. A signal pattern conversion table is built using a radar signal processing simulation system and actual radar experiments, forming an adversarial memory bank for both the radar and the jammer. DDQN samples and trains this memory bank, updating the network parameters. The trained DDQN is then able to determine the optimal jamming scheme.

[0037] This invention constructs a signal pattern conversion table by utilizing error measurement results from radar signal processing systems and experimental data from actual radar systems, thus overcoming the shortcomings of traditional methods and improving the accuracy and applicability of jamming decisions. Furthermore, by introducing the DDQN algorithm and using different value function networks for action selection and evaluation, and implementing action selection and evaluation in the DDQN algorithm with different value function networks, this invention effectively eliminates the problem of Q-value overestimation, improves the accuracy and reliability of jamming decisions, and makes the decision results more effective. This contributes to improving the overall performance and robustness of radar electronic countermeasures systems, meeting the needs of scientific analysis and research, and has positive significance for research on jamming decisions for complex airborne fire control radars. Attached Figure Description

[0038] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0039] Figure 1 This is a flowchart of the method for generating cognitive interference signals for active phased array radar provided by the present invention;

[0040] Figure 2 This is a flowchart of the cognitive interference signal generation algorithm for active phased array radar provided by the present invention;

[0041] Figure 3 This is a schematic diagram of the results of radar signal processing before interference provided by the present invention;

[0042] Figure 4 This is a schematic diagram of radar signal processing after interference provided by the present invention;

[0043] Figure 5 This is a schematic diagram of the DQN algorithm model provided by the present invention;

[0044] Figure 6 This is a schematic diagram of the loss values ​​of the DQN and DDQN algorithms provided by this invention;

[0045] Figure 7 This is a schematic diagram illustrating the maximum Q-value of the DQN and DDQN algorithms provided by this invention;

[0046] Figure 8 This is a schematic diagram of the total reward value of the DQN and DDQN algorithms provided by this invention;

[0047] Figure 9 This is a structural block diagram of the computer device provided by the present invention. Detailed Implementation

[0048] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0049] Figure 1 A schematic diagram of the cognitive jamming signal generation method for active phased array radar according to this embodiment is shown. In this embodiment, the cognitive jamming signal generation method for active phased array radar includes the following steps:

[0050] S1: Construct a radar electronic countermeasures simulation model, and use the TOPSIS sorting algorithm and minimum interference-to-signal ratio data table to obtain the signal pattern state transition table;

[0051] In one exemplary embodiment, step S1 specifically includes:

[0052] S11: Construct a radar electronic countermeasures simulation model, which includes an airborne active phased array radar model and a cognitive jammer model.

[0053] In one exemplary embodiment, the airborne active phased array radar model includes an all-angle-of-entry (AAS) operating mode, a high angle-of-entry (HAS) operating mode, a hybrid angle-of-entry (CAM) operating mode, and a single-target tracking (STT) operating mode. The airborne active phased array radar model also includes a radar signal pattern set, which includes single-target tracking (STT) radar signal patterns, search-plus-track (TAS) radar signal patterns, search-while-track (TWS) radar signal patterns, mid-repetition-frequency search-while-ranging (RWS-1) radar signal patterns, high-repetition-frequency search-while-ranging (RWS-2) radar signal patterns, and velocity-search (VS) radar signal patterns. The cognitive jammer model includes a jamming signal pattern set, which includes comb-pattern jamming signal patterns, sample jamming signal patterns, Doppler noise jamming signal patterns, random false target jamming signal patterns, delay-superimposed agile noise jamming signal patterns, and slice jamming signal patterns.

[0054] As one embodiment, in step S11, a radar electronic countermeasures simulation model is constructed, including constructing an airborne active phased array radar model and a cognitive jammer model; the airborne active phased array radar has multiple operating modes such as AAS, HAS, CAM, and STT, and the N operating modes are denoted as S = {s1, s2, L, s N According to the principle of interference's effect on radar systems, it can be divided into two categories: suppression interference and deception interference. The constructed cognitive jammer model includes noise modulation interference, agile noise interference, comb spectrum interference, Doppler interference, etc., and the M types of interference are denoted as A={a1,a2,L,a... M The signal pattern table 1 shows the signal patterns included in the airborne active phased array radar model.

[0055] Table 1: Signal Pattern Set of Airborne Active Phased Array Radar

[0056]

[0057] Among them, STT is single target tracking, TAS is search plus tracking, TWS is search and track simultaneously, RWS-1 is search and ranging simultaneously at medium repetition rate, RWS-2 is search and ranging simultaneously at high repetition rate, and VS is velocity search; the interference pattern set constructed by the cognitive interference mechanism is shown in Table 2.

[0058] Table 2: Interference Signal Pattern Set

[0059]

[0060] S12: Obtain the error measurement results based on the radar electronic countermeasures simulation model;

[0061] In one exemplary embodiment, the error measurement results include ranging error, speed measurement error, false alarm probability, and detection probability;

[0062] It should be noted that the radar electronic countermeasures simulation model includes a radar signal processing simulation system;

[0063] As one embodiment, in step S12, the radar signal pattern and the jamming signal pattern are input according to the radar electronic countermeasures simulation model; and the corresponding ranging error, velocity error, false alarm probability and detection probability are obtained according to the radar signal processing simulation system.

[0064] Specifically, the simulation program generates radar and interference signals. The radar signal processing system imports the radar and interference signals for signal processing, which includes conventional processes such as pulse compression, Doppler processing, and constant false alarm rate (CFAR) processing, and obtains the final error result. The generated radar and interference signals are then fed into the radar signal processing system. Figure 3 The results of radar signal processing before interference were shown. Figure 3 (a) shows the PD processing result; Figure 3 (b) shows the CFAR processing results; Table 3 shows the radar error measurement results without interference.

[0065] Table 3. Measurement results of error without interference.

[0066]

[0067] The radar signal processing results after adding random decoy interference are as follows: Figure 4 As shown, Figure 4 (a) shows the PD processing result; Figure 4 (b) shows the CFAR processing result.

[0068] Table 4 shows the error measurement results;

[0069] Table 4. Measurement Results of Error Without Interference

[0070]

[0071] As can be seen from the above, the system measures correctly without interference. Then, the interference pattern is selected as dense false target interference, and the interface is displayed as shown above. Through the Doppler response map, it can be seen that multiple false targets are generated. After CFAR processing, multiple targets are still identified, indicating that the interference to the system is successful.

[0072] All interference pattern sets are put into the radar signal processing system for loop traversal to obtain measurement results;

[0073] S13: Use the TOPSIS sorting algorithm to sort the error measurement results and obtain the sorting results;

[0074] In one exemplary embodiment, step S13 specifically includes:

[0075] S131: Based on the error measurement results, the decision matrix is ​​obtained, as shown in the formula:

[0076]

[0077] Where K is the decision matrix, k nm Let be the parameter values ​​corresponding to the nth evaluation index in the mth reconnaissance result;

[0078] S132: Normalize the decision matrix to obtain the normalized matrix, as shown in the formula:

[0079]

[0080] Where, k ij k′ is the element in the i-th row and j-th column of the decision matrix. ij The element in the i-th row and j-th column of the normalized matrix;

[0081] S133: Weight the normalized matrix to obtain the weighted matrix, as shown in the formula:

[0082] z ij =w j k ij ′ (i=1,2,L,m; j=1,2,L,n) (3)

[0083] W = [w1, w2, L, w n ] T ,

[0084] Among them, z ij w is the element in the i-th row and j-th column of the weighted matrix; j Let J be the j-th element of the weight vector, and W be the weight vector.

[0085] S134: Based on the weighted matrix, the optimal reconnaissance result and the worst reconnaissance result are obtained, as shown in the formula:

[0086]

[0087] Among them, I + For optimal reconnaissance results, For the j-th element of the optimal reconnaissance result, I - As the worst reconnaissance result, Let j be the j-th element of the worst reconnaissance result, and max{·} represents taking the maximum value;

[0088] S135: Calculate the distance between each reconnaissance result and the best and worst reconnaissance results respectively, to obtain the first distance and the second distance, as shown in the formula:

[0089]

[0090] in, The first distance, This is the second distance;

[0091] S136: Based on the first and second distances, the quality of each reconnaissance result is determined, as shown in the formula:

[0092]

[0093] Among them, A i The quality of the i-th reconnaissance result;

[0094] S137: The ranking results are obtained based on the quality of each reconnaissance result;

[0095] It should be noted that in step S13, the TOPSIS ranking algorithm is used to rank the error measurement results, that is, to perform decision analysis to approximate the optimal reconnaissance result for multiple evaluation indicators under multiple reconnaissance results; it can be seen from equation (10) that A i The value of A is between 0 and 1; the best reconnaissance result has a quality level of 1, and the worst reconnaissance result has a quality level of 0; therefore, A i The larger the value, the closer the reconnaissance result is to the ideal solution; then, based on the quality of all the reconnaissance results obtained in step 136, the changing trend of the quality of these reconnaissance results will be comprehensively analyzed to obtain the evaluation result of the interference effect; in this embodiment, after radar system simulation experiment, Table 5 shows the ranking result of the TOPSIS of the error measurement results, and VS is set as the final state, and no TOPSIS ranking is performed; the ranking result is shown in Table 5.

[0096] Table 5: Topsis Sorting Results

[0097]

[0098] S14: Use the minimum interference-to-information ratio data table to correct the sorting results and obtain the corrected TOPSIS sorting table;

[0099] As one embodiment, in step S14, the minimum interference-to-signal ratio data table is obtained based on the actual radar interference experiment, and then the TOPSIS ranking results are corrected; wherein, the minimum interference-to-signal ratio test results obtained through the actual radar experiment are shown in Table 6;

[0100] Table 6: Minimum Interference-to-Signal Ratio Test Results (Unit: dB)

[0101]

[0102] In this embodiment, the TOPSIS results are filtered based on the interference-to-signal ratio test results. TOPSIS results of interference patterns with an interference-to-signal ratio of 20dB or higher are removed and will not be considered in the subsequent selection of the optimal interference pattern. The filtered TOPSIS results are shown in Table 7.

[0103] Table 7: TOPSIS Table After Selection

[0104]

[0105] S15: Based on the revised TOPSIS sorting table, obtain the signal pattern state transition table using the transition rules;

[0106] As one embodiment, in step S15, the modified TOPSIS sorting table is converted into a signal pattern state transition table, and the conversion rules are as follows:

[0107] An effective interference pattern against the current signal pattern will change the current signal pattern state to a signal pattern that is relatively ineffective against it. When the interference signal is effective against both RWS-1 and RWS-2, the signal pattern will change to a VS signal, and the VS will be set to a terminated state. Ineffective interference against the current signal pattern will either keep the current signal pattern as it is or change it to a signal pattern of a higher threat level.

[0108] In this embodiment, the signal pattern state transition table is shown in Table 7;

[0109] Table 7 Signal Pattern State Transition Table

[0110]

[0111] It should be noted that effective interference with the current signal pattern will cause the current signal pattern state to change to a signal pattern that is relatively ineffective against it. If there are multiple relatively ineffective signal patterns, a random selection will be made. For example, in the table above, the signal pattern RWS-2 will be switched to signal pattern RWS-2 or TAS by the sampling interference. The signal pattern switching has a certain degree of randomness, which is more in line with the characteristics of phased array radar.

[0112] S2: Based on the radar electronic countermeasures simulation model, the jamming pattern selection model is obtained using the signal pattern state transition table and DDQN network;

[0113] In one exemplary embodiment, step S2 specifically includes:

[0114] S21: Based on the radar electronic countermeasures simulation model, countermeasure samples are obtained using the signal pattern state transition table and DDQN network;

[0115] In one exemplary embodiment, step S21 specifically includes:

[0116] S211: Based on the radar electronic countermeasures simulation model, a set of states and a set of actions are obtained. The set of states includes states, which represent the signal pattern of the enemy's airborne active phased array fire control radar at a certain moment. The set of actions includes actions, which represent the jamming signal pattern selected by our jammer at that moment.

[0117] S212: Based on the state set and action set, use the signal pattern state transition table and DDQN network to obtain the Q value corresponding to each interference signal pattern;

[0118] S213: Based on the Q value corresponding to each interference signal pattern, the first interference signal pattern is obtained using the ε-greed strategy;

[0119] S214: Based on the first interference signal pattern, use the radar electronic countermeasures simulation model and the signal pattern state transition table to obtain the change in the enemy radar signal pattern, and obtain the first report value based on the change in the enemy radar signal pattern.

[0120] S215: Based on the first reward value, update the Q value using the DDQN network, and obtain the next state using the signal pattern state transition table;

[0121] S216: Combine the state, the first interference signal pattern, the first reward value, and the next state into a quadruple sample to obtain the adversarial sample;

[0122] As one embodiment, in step S21, the radar signal pattern and jamming pattern are continuously countered by the radar electronic countermeasures simulation model to generate countermeasure samples;

[0123] It should be noted that traditional Q-learning algorithms have certain limitations when the state and action spaces are large. Q-learning algorithms initialize Q-values ​​and then update them based on state transitions; however, when the state and action spaces are very large, traditional Q-learning algorithms struggle to learn and update effectively. The DQN algorithm model is as follows: Figure 5 As shown; in the DQN algorithm, a neural network is used to approximate the Q-value function; the input state is used as the input of the neural network, and the output is the Q-value of each possible action; by training the neural network, the DQN algorithm can learn the Q-value relationship between the state and the action; the introduction of the DQN algorithm makes it better to process and learn when the state and action space is large; by using a neural network to approximate the Q-value function, the DQN algorithm can handle high-dimensional inputs and can be trained end-to-end through the backpropagation algorithm; the DQN algorithm updates the network parameters θ by setting the number of steps to make the Q function approximate the optimal Q-value, as shown in the following equation (11):

[0124] Q(s,a;θ i )≈Q(s,a)(1)

[0125] In the formula, θ i These are the neural network parameters at the i-th training step.

[0126] DQN has two neural networks with the same structure but different parameters; the main network is used to compute the value function Q(s, a; θ) of the current state-action pair. i The output of the TargetNet is used to compute the Q-estimate. The mathematical expression for calculating the Q estimate is as follows:

[0127]

[0128] In the formula: The parameters representing TargetNet;

[0129] The loss function is as follows:

[0130]

[0131] The parameters of MainNet are updated during each training session, while the parameters of TargetNet are assigned the values ​​of MainNet after a certain number of steps.

[0132] In DQN, the maximization operation when calculating the Q-value may lead to an estimated value function that is higher than its true value, resulting in non-uniform "overestimation," which can affect the final decision. As an offline learning algorithm, DQN does not use the real action of the next interaction during each learning cycle, but instead uses the action currently considered to have the highest value to update the target value function. However, the real policy does not always choose the action with the largest Q-value in a given state, so directly using the maximum Q-value as the target value often leads to the target value being higher than the true value.

[0133] In the DDQN algorithm, action selection and evaluation are implemented through different value function networks. First, the action with the largest Q-value in the next state s' is calculated based on the current network. Then, this action is used to calculate the target Q-value, and the target network selects the optimal action. The calculation method of the target Q-value in DDQN is as follows:

[0134]

[0135] Substituting equation (14) into the loss function calculation formula, we get:

[0136] Loss(θ)=E[(TargetQ-Q(s,a;θ i )) 2 (5)

[0137] In one exemplary embodiment, the principle of the DDQN algorithm is applied to the decision-making of jamming signal patterns. The following variables are defined: state (s∈S) represents the signal pattern of the enemy's airborne active phased array fire control radar at a certain moment, and set (S) represents all possible enemy signal patterns; action (a∈A) represents the jamming signal pattern selected by our jammer at that moment, and set (A) is the set of jamming signal patterns available to our side; after our reconnaissance equipment determines the state (s) through radar signal pattern identification, it inputs it into the DDQN network; the network calculates the Q value corresponding to each jamming signal pattern through neural network fitting; and according to strategy ε-g... The reed strategy (used to balance exploration and exploitation to maximize cumulative reward) selects a jamming signal pattern (a) to apply to the enemy radar; after jamming, our side evaluates the effect based on the change in the enemy radar signal pattern and obtains a reward (r∈R); since our side and the enemy are in a non-cooperative relationship, our side can evaluate the effectiveness of jamming by observing the change in the signal pattern after jamming, and thus give the corresponding reward value (r); the jammer uses the reward value to update the Q value to determine the jamming pattern to be adopted in the next state (s'), and stores the above four variables into a quadruple sample (s, a, r, s') in the sample pool (D);

[0138] S22: Construct and update the memory based on adversarial examples;

[0139] As one embodiment, in step S22, an adversarial memory bank is constructed based on the sample pool (D), and adversarial samples are stored in the memory bank. When the number of samples exceeds the capacity of the memory bank, new samples will replace old samples.

[0140] S23: Using the memory bank as training samples, train the DDQN network to obtain the interference pattern selection model;

[0141] As one embodiment, in step S23, DDQN randomly selects samples from the memory bank as training samples and continuously updates the network parameters during the training process; after reaching the total number of iterations, a trained interference pattern selection model is obtained; specifically, in the sample pool (D), a certain number of samples are randomly selected for neural network training to update the network parameters; then, the interference signal pattern selection process is repeated until the termination state is reached.

[0142] It should be noted that since different interference styles produce different interference effects, the reward value is defined as follows:

[0143]

[0144] In the formula: JSR represents the minimum interference-to-signal ratio; TL→min indicates that the signal pattern threat level of the airborne active phased array fire control radar will reach the lowest level; TL↓ indicates that the signal pattern will change towards a lower threat level; TL=TL indicates that there is no change between signal patterns; TL↑ indicates that the signal pattern will change towards a higher threat level.

[0145] S3: Use the interference pattern selection model to make the optimal interference pattern decision.

[0146] In one exemplary embodiment, a simulation experiment was conducted using the aforementioned active phased array radar cognitive jamming signal generation method. Figure 2 The diagram shows the flowchart of the cognitive jamming signal generation algorithm for an active phased array radar. Details are as follows:

[0147] In this embodiment, different threat levels are set for different signal patterns, which are used as the basis for evaluating the interference effect. STT is the highest level, set to 4, TAS is 3, TWS is 2, RWS-1 and RWS-2 are set to 1, and VS is set to 1. The signal pattern conversion table is shown in 3.2.1. The algorithm training parameters in the simulation experiment are shown in Table 8.

[0148] Table 8: Training Parameters

[0149] parameter Value Discount factor 0.9 Learning rate 0.005 Memory capacity 500 Training activation threshold 200 Sampling size 64 Network replication frequency 5 Greed factor [0.1-0.95]

[0150] Before exchanging parameters between the two neural networks, we first observe for 200 steps; then, every 5 steps, we assign the estimated network parameters to the target network; in the greedy selection, the initial value is set to 0.1 and gradually increased by 0.0005 per step until the termination value of 0.95 is reached; each time, 64 sets of samples are randomly drawn from the memory bank for network training, and the network parameters are continuously updated through the backpropagation algorithm to improve the estimation accuracy of the Q value;

[0151] Figure 6 These are the loss maps obtained from training the DQN and DDQN algorithms. Due to the low ε exploration value at the initial moment, randomly selecting interference patterns may cause drastic fluctuations in the loss value. However, as training progresses, the ε exploration value gradually increases, and the neural network's fit to the Q value becomes more accurate. After approximately 300 steps, the loss values ​​of both algorithms drop to near 0. Although the ε exploration value increases to a maximum of 0.95, random selection of interference signal patterns for exploration still occurs in the later stages of training, so the loss values ​​of both algorithms also show slight fluctuations in the later stages of training.

[0152] The training process records the maximum Q-value of a certain state output by the current network, such as... Figure 7 As shown; Figure 7 The maximum Q-value convergence of the two algorithms at RWS-M is shown. The DQN algorithm converges to around 100 after 500 iterations, while the DDQN algorithm converges to around 100 after 400 iterations. Comparing the maximum Q-value convergence of the two algorithms, it can be seen that the DDQN converges faster than the DQN and eliminates some of the overestimation of Q-values.

[0153] The ultimate goal of reinforcement learning algorithms is to maximize the total objective reward. Therefore, the total reward is obtained by summing the reward values ​​generated by each perturbation pattern selection during training, such as... Figure 8 As shown, after 100 training steps, the total reward value obtained by the DDQN algorithm is higher than that of the DQN algorithm. This indicates that the DDQN algorithm is superior to the DQN algorithm in the selection of interference patterns, thus verifying the effectiveness of the proposed method.

[0154] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method for generating cognitive jamming signals for an active phased array radar. The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium may also include combinations of the above types of memory.

[0155] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the above-described method for generating cognitive jamming signals for active phased array radar.

[0156] like Figure 9 As shown, the computer device 120 may include: at least one processor 121, such as a central processing unit (CPU), at least one communication interface 123, memory 124, and at least one communication bus 122. The communication bus 122 is used to enable communication between these components. The communication interface 123 may include a display screen and a keyboard; optionally, the communication interface 123 may also include a standard wired interface or a wireless interface. The memory 124 may be high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory 124 may also be at least one storage device located remotely from the aforementioned processor 121. The memory 124 stores application programs, and the processor 121 calls the program code stored in the memory 124 to execute any of the aforementioned method steps. The communication bus 122 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus 122 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 9The bus is represented by a single line, but this does not imply a single bus or a single type of bus. The memory 124 may include volatile memory, such as random-access memory (RAM); it may also include non-volatile memory, such as flash memory, hard disk drive (HDD), or solid-state drive (SSD); or it may include combinations of the above types of memory. The processor 121 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP. The processor 121 may further include hardware chips. These hardware chips may be application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or combinations thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. Optionally, the memory 124 is also used to store program instructions. The processor 121 can call the program instructions to implement the active phased array radar cognitive jamming signal generation method as described in this embodiment.

[0157] This embodiment provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method for generating cognitive jamming signals for active phased array radar.

[0158] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method for generating cognitive jamming signals for an active phased array radar, characterized in that, Includes the following steps: S1: Construct a radar electronic countermeasures simulation model, and use the TOPSIS sorting algorithm and minimum interference-to-signal ratio data table to obtain the signal pattern state transition table; S2: Based on the radar electronic countermeasures simulation model, the jamming pattern selection model is obtained using the signal pattern state transition table and the DDQN network; S3: Utilize the aforementioned interference pattern selection model to make the optimal interference pattern decision; Step S1 specifically includes: S11: Construct a radar electronic countermeasures simulation model, which includes an airborne active phased array radar model and a cognitive jammer model; S12: Obtain the error measurement results based on the radar electronic countermeasures simulation model; S13: Sort the error measurement results using the TOPSIS sorting algorithm to obtain the sorting results; S14: The sorting results are corrected using the minimum interference-to-information ratio data table to obtain the corrected TOPSIS sorting table; S15: Based on the modified TOPSIS sorting table, obtain the signal pattern state transition table using the transformation rules; Step S2 specifically includes: S21: Based on the radar electronic countermeasures simulation model, use the signal pattern state transition table and DDQN network to obtain the countermeasures samples; S22: Construct and update the memory bank based on the adversarial examples; S23: Using the memory bank as training samples, train the DDQN network to obtain the interference pattern selection model; Step S21 specifically includes: S211: Based on the radar electronic countermeasures simulation model, a state set and an action set are obtained; the state set includes states, which represent the signal pattern of the enemy's airborne active phased array fire control radar at a certain moment; the action set includes actions, which represent the jamming signal pattern selected by our jammer at that moment. S212: Based on the state set and action set, the Q value corresponding to each interference signal pattern is obtained using the signal pattern state transition table and the DDQN network; S213: Based on the Q value corresponding to each interference signal pattern, using... The strategy yields the first interference signal pattern; S214: Based on the first interference signal pattern, the change in the enemy radar signal pattern is obtained using the radar electronic countermeasures simulation model and the signal pattern state transition table, and the first report value is obtained based on the change in the enemy radar signal pattern. S215: Based on the first return value, update the Q value using the DDQN network, and obtain the next state using the signal pattern state transition table; S216: Combine the state, the first interference signal pattern, the first reward value, and the next state into a quadruple sample to obtain an adversarial sample.

2. The method for generating cognitive jamming signals for active phased array radar according to claim 1, characterized in that, The airborne active phased array radar model includes the full angle of entry search (AAS) working mode, the high angle of entry search (HAS) working mode, the hybrid angle of entry search (CAM) working mode, and the single target tracking (STT) working mode. The airborne active phased array radar model also includes a radar signal pattern set, which includes single target tracking (STT) radar signal pattern, search and track (TAS) radar signal pattern, search and track simultaneously (TWS) radar signal pattern, medium repetition frequency search and ranging simultaneously (RWS-1) radar signal pattern, high repetition frequency search and ranging simultaneously (RWS-2) radar signal pattern, and velocity search (VS) radar signal pattern. The cognitive jamming machine model includes a set of jamming signal patterns; the set of jamming signal patterns includes a comb pattern jamming signal pattern, a sample jamming signal pattern, a Doppler noise jamming signal pattern, a random false target jamming signal pattern, a delayed superposition clever noise jamming signal pattern, and a slice jamming signal pattern.

3. The method for generating cognitive interference signals for active phased array radar according to claim 1, characterized in that, The error measurement results include ranging error, speed measurement error, false alarm probability, and detection probability.

4. The method for generating cognitive interference signals for active phased array radar according to claim 1, characterized in that, Step S13 specifically includes: S131: Based on the error measurement results, the decision matrix is ​​obtained, as shown in the formula: , in, For the decision matrix, For the first The evaluation indicators in the first The parameter values ​​corresponding to each reconnaissance result; S132: Normalize the decision matrix to obtain the normalized matrix, as shown in the formula: , in, The decision matrix is ​​the first Line 1 Column elements; For the normalized matrix, the first... Line 1 Column elements; S133: Weight the normalized matrix to obtain the weighted matrix, as shown in the formula: , , in, The weighted matrix is ​​the first Line 1 Column elements; The weight vector of the first One element, This is the weight vector; S134: Based on the weighted matrix, the optimal reconnaissance result and the worst reconnaissance result are obtained, as shown in the formula: , , , , in, For optimal reconnaissance results, The first optimal reconnaissance result One element, As the worst reconnaissance result, The worst reconnaissance result One element, This indicates taking the maximum value; S135: Calculate the distance between each reconnaissance result and the best and worst reconnaissance results respectively, to obtain the first distance and the second distance, as shown in the formula: , , in, The first distance, This is the second distance; S136: Based on the first and second distances, the quality of each reconnaissance result is determined, as shown in the formula: , in, For the first The quality of each reconnaissance result; S137: Based on the quality of each reconnaissance result, a ranking result is obtained.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the active phased array radar cognitive jamming signal generation method as described in any one of claims 1-4.

6. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the active phased array radar cognitive jamming signal generation method as described in any one of claims 1-4.

7. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the active phased array radar cognitive jamming signal generation method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Distribution network investment decision analysis model based on improved Topsis method

    CN109034511A

  • Interference efficiency library construction method and system for unknown radar signals

    CN116151004A

  • Constant-interference-to-signal-ratio active stealth target interference method and system

    CN116794611A