Radar interference signal generation method based on curiosity driven Q learning and LSTM prediction

By employing a curiosity-driven Q-learning and LSTM prediction method, the adaptability and real-time performance issues of traditional radar jamming waveforms in complex electromagnetic environments are addressed. This generates highly adaptable and real-time jamming signals, enabling effective radar jamming in complex environments.

CN121856909APending Publication Date: 2026-04-14CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Traditional radar jamming waveforms are difficult to adapt to real-time changes in complex electromagnetic environments, resulting in high computational overhead, poor real-time performance, weak environmental adaptability, and limited jamming effect.

Method used

A curiosity-driven Q-learning and LSTM prediction approach is adopted. The interference waveform is optimized through a Rayleigh entropy-driven reward and punishment mechanism. By combining transfer learning and LSTM network, the curiosity mechanism and dual prediction scheme are dynamically triggered to generate interference signals with strong adaptability and high real-time performance.

Benefits of technology

It achieves real-time performance and improved effectiveness of radar jamming in complex electromagnetic environments, and can completely mask the original signal in scenarios with a signal-to-noise ratio of 0dB and multiple types of radar signals, with robustness superior to traditional technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121856909A_ABST
    Figure CN121856909A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of radar interference, and particularly discloses a radar interference signal generation method based on curiosity driven Q learning and LSTM prediction, and the method comprises the steps: selecting an initial interference waveform parameter from a simulation waveform library based on Q learning; in the Q learning process, monitoring the Rayleigh entropy change of the interference signal of each iteration, when the difference between the Rayleigh entropy of the current step and the Rayleigh entropy of the previous step exceeds a threshold value, exploring a new action outside the waveform library in an error range, and screening the interference signal with the highest Rayleigh entropy; the frozen Q value table is used as an initial value, formula fine tuning parameters are updated according to the Q value, and an interference waveform adaptive to a new environment is generated; and inputting the obtained interference waveform into the long short-term memory network for prediction optimization, generating a predicted interference signal through two schemes of gradually predicting partial data by the training set and performing one-step full-quantity prediction by the test set, and selecting the interference signal with higher Rayleigh entropy as a final optimized interference signal. According to the invention, the radar interference effect, the real-time performance and the environmental adaptability can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of radar jamming technology, and more specifically, relates to a radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction. Background Technology

[0002] As a core detection device in electronic warfare systems, radar's ability to detect targets and acquire information directly impacts the effectiveness of combat decision-making. Radar jamming, by generating specific interference waveforms to suppress or deceive radar, prevents it from extracting effective information such as target position, distance, and speed, and is a key means of electronic countermeasures. With the rapid development of cognitive electronic warfare and artificial intelligence technologies, the battlefield electromagnetic environment is characterized by "high density, dynamism, and complexity," making it difficult for traditional fixed-parameter interference waveforms to match the real-time changes in radar signals.

[0003] Therefore, there is an urgent need for an interference waveform generation technology with real-time learning, adaptive optimization, and environmental adaptability to meet the radar countermeasure requirements in complex electromagnetic environments. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this application aims to provide a radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction, which can effectively improve radar jamming performance, real-time performance, and environmental adaptability, thereby enhancing the ability to suppress radar jamming in complex electromagnetic environments.

[0005] To achieve the above objectives, in a first aspect, this application provides a radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction, comprising the following steps: S10: Select initial interference waveform parameters from the simulated waveform library based on Q-learning. Rayleigh entropy is used as instantaneous reward and punishment feedback for Q-learning. The action selection is carried out iteratively to update the Q-value table through a greedy strategy to obtain the initial optimal interference waveform parameters. S20: During Q-learning, monitor the change of Rayleigh entropy value of the interference signal in each iteration. When the difference between the Rayleigh entropy value of the current step and the Rayleigh entropy value of the previous step exceeds the threshold, it is judged as an error. Within the error range, explore new actions outside the waveform library and screen the interference signal with the highest Rayleigh entropy to optimize waveform selection performance. S30: Freeze the trained Q-value table and migrate it to the new electromagnetic environment. Using the frozen Q-value table as the initial value, select actions through a greedy strategy and fine-tune parameters according to the Q-value update formula to generate interference waveforms adapted to the new environment. S40, the interference waveforms obtained in steps S20 and S30 are input into the long short-term memory network for prediction optimization. The predicted interference signal is generated by two methods: gradually predicting part of the data in the training set and predicting the full data in one step in the test set. The predicted interference signal is then superimposed with the interference waveforms obtained in step S30, and the signal with the higher Rayleigh entropy is selected as the final optimized interference signal.

[0006] The radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction provided in this application has the following advantages: First, the Rayleigh entropy-driven reward and punishment mechanism quantifies the jamming effect into an iteratively optimizable index, avoiding the high computational overhead of traditional criterion functions traversing waveform libraries. Second, the curiosity-driven mechanism's dynamic triggering strategy only initiates additional searches when significant errors occur during exploration, balancing exploration depth and efficiency. Third, the synergy between transfer learning and LSTM allows the model to adapt to new environments by freezing the Q-table, and further optimizes the waveform by combining it with a dual LSTM prediction scheme, achieving a dual improvement in real-time performance and jamming effect. Furthermore, this method can completely mask the original signal even in scenarios with a signal-to-noise ratio of 0dB and various types of radar signals, demonstrating superior robustness compared to traditional techniques. Moreover, the core algorithm modules (curiosity-driven Q-learning and LSTM optimization) have clear principles and can be extended to scenarios such as communication signal jamming and satellite anti-jamming according to actual needs, exhibiting strong adaptability.

[0007] As a further preferred option, step S10 specifically involves: constructing an analog waveform library using linear frequency modulated signals and square waves; using Rayleigh entropy as the instantaneous reward and punishment feedback basis for Q-learning; combining generalized linear frequency modulated transform and time redistribution multi-synchronous compression transform for time-frequency analysis; and using a greedy strategy to select actions to iteratively update the Q-value table to obtain the initial optimal interference waveform parameters.

[0008] As a further preferred embodiment, the linear frequency modulation signal is represented as:

[0009] Square waves are represented as:

[0010] In the formula, FA is the square wave amplitude, and DR is the duty cycle. Let K be the carrier frequency and K be the linear modulation frequency of the signal. The form of the interference signal output by this waveform library is: .

[0011] As a further preferred embodiment, in step S20, it is assumed that the Rayleigh entropy value of the interference waveform detected in the previous step is... ( i>1 The Rayleigh entropy value of the interference signal detected in the current step is... When satisfied At that time, the curiosity mechanism is triggered.

[0012] As a further preferred embodiment, in step S30, the Q-value update formula used for fine-tuning the parameters is:

[0013] In the formula, In order to be in t The state of the surrounding environment at any given moment; In order to be in t At this moment, the environmental state is The action performed at that time; The learning rate represents the next state; Indicates the learning rate; Indicates the discount rate; r This provides instantaneous feedback for each action.

[0014] As a further preferred embodiment, in step S40, when the Long Short-Term Memory Network performs prediction optimization, the input data is divided into a training set and a test set in an 8:2 ratio. The training set is used to gradually predict 20% of the data to generate a first prediction interference signal, and the test set is used to predict the entire dataset in one step to generate a second prediction interference signal. The Rayleigh entropy of the first prediction interference signal and the second prediction interference signal is calculated, and the signal with the higher Rayleigh entropy is selected as the final interference signal.

[0015] As a further preferred embodiment, in step S40, the long short-term memory network includes a forget gate, an input gate, and an output gate. The forget gate controls the retention or forgetting of information in the cell state, the input gate controls the updating of the cell state by the current input information, and the output gate controls the influence of the cell state on the hidden state.

[0016] As a further preferred embodiment, in step S40, the operation of the forget gate is as follows: [The text abruptly ends here, likely due to an incomplete sentence or a formatting error.] t Hidden state information at time 1 With the current moment t Data Common input The function outputs a value between 0 and 1, and then... The value output by the function is the same as the memory from the previous moment. To limit the influence of memories from a previous moment on memories from subsequent moments by multiplying them, the formula is as follows:

[0017] In the formula, It is a measure of the importance of past memories. It is the input at time t. It is the data from the previous hidden layer. It is used to Adjust the matrix to have the same dimensions as the hidden layer at time t. For bias.

[0018] As a further preferred embodiment, the operation process of the input gate is as follows: First Enter information at any time Hidden information from the previous moment After a The function will then the current Enter information at any time Hidden information from the previous moment After passing through a tanh layer, a new candidate value vector is created; then the two values ​​are multiplied to determine the amount of input information retained at the current time step.

[0019] Secondly, this application provides an application of the method as described in any one of the above statements, applicable to the fields of military electronic warfare or civilian radar.

[0020] It is understandable that the beneficial effects of the second aspect mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0021] Figure 1 This is a flowchart of the radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction provided in this application; Figure 2 This is a technical block diagram of the radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction provided in the embodiments of this application; Figure 3 This is a schematic diagram of reinforcement learning provided in an embodiment of this application; Figure 4 This is a diagram illustrating the role of curiosity in an exploratory environment, provided in an embodiment of this application. Figure 5 This is a diagram illustrating the structure of a Long Short-Term Memory (LSTM) network provided in an embodiment of this application. Figure 6 This is a simulation step diagram of the Q-learning algorithm provided in the embodiments of this application; Figure 7 The results of Q-learning input signal and output interference signal provided in the embodiments of this application are shown; where (a) is the interference signal, (b) is the time domain diagram, (c) is the frequency domain diagram, and (d) is the time-frequency diagram. Figure 8 These are comparison diagrams showing the results after executing the curiosity mechanism provided in the embodiments of this application; where (a) is a signal, (b) is a time-domain diagram, (c) is a frequency-domain diagram, and (d) is a time-frequency diagram; Figure 9These are comparison diagrams before and after interference obtained after transfer learning according to the embodiments of this application; wherein, (a) is the interference signal, (b) is the time domain diagram, (c) is the frequency domain diagram, and (d) is the time-frequency diagram; Figure 10 These are comparison diagrams of optimized waveforms based on transfer learning provided in the embodiments of this application; wherein, (a) is the best interference signal, (b) is the time domain diagram before and after interference, (c) is the frequency domain diagram before and after interference, and (d) is the time-frequency diagram; Figure 11 This is a comparison between the CDQL-LSTM algorithm and the CDM-TL algorithm provided in the embodiments of this application; wherein, (a) is LSTM 20% prediction, (b) is LSTM 100% prediction, (c) is the interference result, and (d) is the frequency domain diagram; Figure 12 The images provided in this application are before and after interference based on the CDQL-LSTM algorithm; where (a) is the interference signal, (b) is the time domain diagram, (c) is the frequency domain diagram, and (d) is the time-frequency diagram. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0023] like Figure 1 As shown, this application provides a radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction, including steps S10 to S40, which are detailed below: Step S10: Select initial interference waveform parameters from the simulated waveform library based on Q-learning. Rayleigh entropy is used as instantaneous reward and punishment feedback for Q-learning. The Q-value table is updated iteratively by selecting actions through a greedy strategy to obtain the initial optimal interference waveform parameters.

[0024] This step uses Rayleigh entropy quantization of the interference effect as a reward or punishment feedback, which can guide the agent to quickly select the optimal interference parameters, reduce computational overhead, and improve the efficiency of initial waveform generation.

[0025] Step S20: During Q-learning, monitor the Rayleigh entropy value change of the interference signal in each iteration. If the difference between the Rayleigh entropy value of the current step and the Rayleigh entropy value of the previous step exceeds the threshold, it is judged as an error. Explore new actions outside the waveform library within the error range, and filter the interference signal with the highest Rayleigh entropy to optimize the waveform selection performance.

[0026] Step S20 uses a dynamic curiosity mechanism to explore potential better parameters outside the waveform library when a significant error is detected. This avoids ineffective exploration and improves the optimization accuracy and convergence speed of the interference waveform.

[0027] Step S30: Freeze the trained Q-value table and migrate it to the new electromagnetic environment. Using the frozen Q-value table as the initial value, select actions through a greedy strategy and fine-tune parameters according to the Q-value update formula to generate an interference waveform adapted to the new environment.

[0028] Step S30 utilizes transfer learning to quickly adapt existing knowledge to the new environment. By fine-tuning parameters, retraining time can be reduced, thereby improving the algorithm's adaptability to dynamic electromagnetic environments and its real-time response capability.

[0029] In step S40, the interference waveforms obtained in steps S20 and S30 are input into the long short-term memory network for prediction optimization. The predicted interference signal is generated by two methods: gradually predicting part of the data in the training set and predicting the full data in one step in the test set. The predicted interference signal is then superimposed with the interference waveform obtained in step S30, and the signal with the higher Rayleigh entropy is selected as the final optimized interference signal.

[0030] Step S40 utilizes the temporal modeling capability of long short-term memory networks to generate a better interference signal through a dual prediction scheme, and combines it with a superposition selection mechanism to further improve the quality and robustness of the interference waveform.

[0031] The radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction provided in this application has the following advantages: First, the Rayleigh entropy-driven reward and punishment mechanism quantifies the jamming effect into an iteratively optimizable index, avoiding the high computational overhead of traditional criterion functions traversing waveform libraries. Second, the curiosity-driven mechanism's dynamic triggering strategy only initiates additional searches when significant errors occur during exploration, balancing exploration depth and efficiency. Third, the synergy between transfer learning and LSTM allows the model to adapt to new environments by freezing the Q-table, and further optimizes the waveform by combining it with a dual LSTM prediction scheme, achieving a dual improvement in real-time performance and jamming effect. Furthermore, this method can completely mask the original signal even in scenarios with a signal-to-noise ratio of 0dB and various types of radar signals, demonstrating superior robustness compared to traditional techniques. Moreover, the core algorithm modules (curiosity-driven Q-learning and LSTM optimization) have clear principles and can be extended to scenarios such as communication signal jamming and satellite anti-jamming according to actual needs, exhibiting strong adaptability.

[0032] Specifically, the method provided in this application can be applied to the fields of radar jamming signal generation, electromagnetic environment adaptation, and anti-jamming performance testing. In the field of military electronic warfare, it can be integrated into radar countermeasures equipment for fighter jets, warships, and electronic warfare aircraft to generate suppressive / deceptive jamming signals, interfering with enemy early warning radars and fire control radars, and preventing them from obtaining key information such as target position and speed. Simultaneously, it can serve as a core module of military electronic countermeasures training systems, simulating various types of jamming scenarios in complex electromagnetic environments, thereby enhancing the response capabilities of combat personnel. In the field of civilian radar, it can be adapted to air traffic control radars, weather radars, and maritime radars to generate targeted jamming signals to counteract illegal electromagnetic noise or complex environmental interference, ensuring flight safety, meteorological data accuracy, and the reliability of maritime target identification. The improved transfer learning and curiosity-driven mechanism strategy can also optimize the signal generation process for other electromagnetic interference scenarios, further expanding the application scope.

[0033] In one embodiment, the technical solution to achieve the above objective can be as follows: To address the problems of high computational overhead, poor real-time performance, and weak adaptability to unknown electromagnetic environments in current radar jamming signal generation technologies, this embodiment provides a radar jamming signal optimization method based on curiosity-driven Q-learning (CDQL) and long short-term memory (LSTM) networks combined with transfer learning. This method utilizes a multi-mechanism collaboration of "curiosity-driven Q-Learning optimization strategy generation - transfer learning to accelerate environment adaptation - LSTM optimization of time-domain waveforms." Through the collaborative framework of curiosity-driven Q-learning + LSTM prediction, multi-dimensional technical breakthroughs are achieved. Rayleigh entropy quantifies the reward and penalty feedback of Q-learning, and generalized linear frequency modulation transform and time redistribution multi-synchronous compression transform improve the learning efficiency of jamming parameters. Simultaneously, a curiosity-driven mechanism (CDM) dynamically triggers and optimizes ineffective exploration, balancing optimization accuracy and efficiency, enabling rapid model adaptation to new environments. This solves the problems of poor real-time performance, weak environmental adaptability, and limited jamming effect of traditional jamming waveform generation methods, effectively improving the suppression capability of radar jamming in complex electromagnetic environments. The implementation idea is as follows: Figure 2 As shown.

[0034] Specifically, the radar interference signal generation method provided in this embodiment is as follows: Step 1: Q-learning initial interference parameter learning: Construct a simulated waveform library using linear frequency modulated (LFM) signals and square waves, and use Rayleigh entropy as the instantaneous reward and punishment feedback basis for Q-learning. Combine generalized linear frequency modulated transform and time redistribution multi-synchronous compression transform for time-frequency analysis; select actions through a greedy strategy, iteratively update the Q table, obtain the initial optimal interference waveform parameters, and realize the optimization of interference parameters for known radar signals.

[0035] Step 2: Curiosity-Driven Mechanism (CDM) Optimization: Monitor the Rayleigh entropy difference of the interference signal in each iteration. When the difference exceeds 5% of the entropy value of the previous step, CDM is triggered. Based on Q-learning exploration, explore new actions outside the original waveform library within the error range, store and filter the interference signal with the highest Rayleigh entropy, and optimize waveform selection performance.

[0036] Step 3: Transfer learning to adapt to the new environment: Freeze the Q-table of the Q-learning-curiosity mechanism model after training and transfer it to the new electromagnetic environment; using the frozen Q-table as the initial value, select actions through a greedy strategy, fine-tune parameters according to the Q-value update formula, accelerate model convergence in the new environment, and generate interference waveforms adapted to the new scene.

[0037] Step 4: LSTM Network Waveform Prediction Optimization: Using the waveform obtained by combining transfer learning with CDM as input, the training set and test set are divided in an 8:2 ratio. An LSTM network with 300 hidden layers is constructed and trained for 500 rounds. Two schemes are adopted: "gradually predicting 20% ​​of the data in the training set" and "one-step full prediction in the test set". After superimposing the predicted waveform with the transfer learning waveform, the waveform with the higher Rayleigh entropy is selected as the final optimized interference waveform.

[0038] (a) Curiosity-driven Q-learning (1) Q learning It should be noted that Q-learning (QL) is a model-free reinforcement learning method. Q-learning typically includes a state set, an action set, and an instantaneous reward / punishment feedback R. The principle diagram of reinforcement learning is as follows. Figure 3 As shown.

[0039] In interactive learning, the agent does not know the optimal policy and relies solely on the instantaneous reward and punishment feedback from the environment to calculate the reward value, aiming to maximize the gain. The policy π is defined as the mapping from state S to action A. Let be the state of the surrounding environment at time t. Let the environmental state be at time t. The action performed at that time. The value of policy π is derived from the environmental state. Let r be the total reward obtained after performing N actions, where r is the instantaneous feedback for each action.

[0040] (1) Equation (1) represents the value of policy π obtained by performing N actions in a finite-order model. However, in an infinite-order model, where the sequence is infinitely long, future rewards will be discounted, and the value of policy π in this case is: (2) in This indicates the size of the discount rate.

[0041] For each policy chosen by the agent, the goal is to maximize the value of π. This embodiment defines the optimal policy in each policy as... The value of the optimal strategy is shown in equation (3): (3) In Q-learning, the value function of policy π is replaced by the state-action value, denoted as . It indicates that at time t, the environmental state is When the intelligent agent executes an action The value obtained. The basic form of Q-value update is shown in equation (4): (4) in, This indicates that the agent is in a state. Below, take action The optimal reward discount obtained, This represents the learning rate, which gradually decreases as the learning process progresses. This indicates the discount rate. A larger value indicates that future rewards are more important than current rewards. In the Q-learning algorithm, actions are selected using a greedy strategy. That is, in terms of probability Randomly select an action Otherwise, the strategy corresponding to the maximum value function is selected, i.e., according to... Corresponding To select the current state The following action .

[0042] Due to the increasingly complex electromagnetic environment, this embodiment combines the Generalized Linear Frequency Modulation (GLCT) transform with the Time Reassignment Multiple Synchronous Compression (GSMC) transform. The GLCT transform is based on the Linear Frequency Modulation (LCT) transform. It stores the time-frequency representations obtained from N different frequency LCT transforms into a three-dimensional matrix, and then extracts the portion of each LCT with the highest energy concentration from the three-dimensional matrix to form the GLCT. The Time Reassignment Multiple Synchronous Compression (GSMC) transform first performs an STFT on the signal, according to equation (5) (where...). The initial two-dimensional group delay is calculated from the result of the short-time Fourier transform of the signal. Then, fixed-point iteration is performed on the initial two-dimensional group delay to obtain a result that approximates the true group delay (i.e., the signal phase). The two-dimensional group delay is estimated, and then the time-frequency representation obtained by the short-time Fourier transform (STFT) is frequency compressed in the time direction so that the energy is compressed into the group delay trajectory. Then the compressed result is iterated again, which is the redistribution process.

[0043] (5) In this embodiment, the simulation uses a linear frequency modulated (LFM) signal as the interference waveform, and then different parameters are combined to form different waveforms, constituting the waveform library for simulation. In the waveform library constructed in this embodiment, the LFM signal is represented as shown in the following equation (6): (6) The square wave is represented as shown in equation (7): (7) Where FA is the square wave amplitude and DR is the duty cycle. The final output interference signal takes the form of... .

[0044] (2) Curiosity Mechanism This embodiment uses Q-learning to select waveforms from a waveform library. However, in reality, due to the lack of prior knowledge, the exploration performed by Q-learning may involve ineffective exploration. Therefore, a curiosity-driven mechanism (CM) is proposed to optimize the results of Q-learning. Curiosity is a mechanism that drives an agent to learn new knowledge and skills by motivating the agent to take action, thereby reducing its uncertainty about behavior and consequences. The agent receives higher intrinsic motivation when the error between prediction and actual behavior is large, thus continuously exploring, interacting, and learning.

[0045] like Figure 4 As shown, under sparse external rewards, curiosity can reduce environmental interaction to achieve goals; without external rewards, curiosity drives the agent to explore efficiently; facing unknown environments, early experience can help the agent explore new areas. This embodiment uses an error-triggered curiosity mechanism to optimize the exploration of interference waveforms in Q-learning. First, the Rayleigh entropy value of the interference signal in each iteration is stored in a matrix. The initial exploration is driven only by external stimuli, while subsequent explorations are based on changes in Rayleigh entropy value that trigger the curiosity mechanism: if the difference between the Rayleigh entropy value of the current step and the previous step exceeds a threshold, it is judged as an error, prompting the agent to explore new actions within this error range. These actions are not in the original waveform library. Finally, based on Q-learning, a potentially better interference waveform is generated through the curiosity mechanism.

[0046] Assume the Rayleigh entropy value of the interference waveform obtained in the previous exploration is (i>1), the Rayleigh entropy of the signal obtained in this exploration is... When equation (8) is satisfied, the curiosity mechanism is triggered.

[0047] (8) (3) Transfer learning It should be noted that transfer learning (TL) is the application of knowledge or patterns learned in one domain or task to different but related domains or problems. It is a method based on the similarity of data, tasks, and models, which transfers knowledge learned in one domain to another similar domain.

[0048] Long Short-Time Memory (LSTM) networks are commonly used for predicting time-series signals. The network structure of an LSTM is as follows: Figure 5 As shown, it consists of three parts: a forget gate, an input gate, and an output gate. The forget gate controls which information in the cell state should be retained or forgotten to optimize long-term memory; the input gate controls the degree to which the current input information updates the cell state, determining which information should be remembered; the output gate controls the influence of the current cell state on the hidden state, determining the output information to be passed to the next layer.

[0049] (4) Q-learning algorithm test The simulation steps of the Q-learning algorithm provided in this embodiment are as follows: Figure 6 As shown.

[0050] The LFM signal is represented as shown in equation (9): (9) The square wave is represented as shown in equation (10): (10) The final output interference signal is in the form of The following is the final output of Q-learning, such as... Figure 7 As shown.

[0051] An interference waveform can be obtained using Q-learning, from which... Figure 7 As shown in (d), the original input signal is completely covered after interference, indicating that the interference signal has a significant masking effect. According to the data in Table 1, the Rayleigh entropy value increases significantly after interference, indicating that the time-frequency concentration of the signal decreases and the amount of information is reduced, confirming that the interference achieved the expected effect.

[0052] Table 1 Rayleigh entropy values ​​before and after Q-learning interference

[0053] (5) Q-learning simulation based on curiosity mechanism The curiosity mechanism triggering simulation steps provided in this embodiment are as follows: First, it is determined whether the curiosity mechanism is triggered. If triggered, the curiosity model is entered, and the Rayleigh entropy value of the signal is stored and calculated. Next, Q-learning iterations are performed until completion. During the iteration process, the Rayleigh entropy values ​​of the curiosity mechanism and Q-learning results are calculated and compared, and the signal corresponding to the maximum value is selected. Finally, the signal obtained by the curiosity mechanism is compared with the Q-learning result. If the Rayleigh entropy value of the interference signal obtained by the curiosity mechanism is larger, the result of the curiosity mechanism is output; otherwise, the result of Q-learning is output. In 5000 Q-learning iterations, the curiosity mechanism is triggered 1095 times. The optimal interference result obtained by executing the curiosity mechanism is as follows: Figure 8 As shown: Table 2 Rayleigh entropy values ​​before and after interference with the Q-learning-curiosity mechanism

[0054] The above results show that the solution obtained by the curiosity mechanism is better than that obtained by Q-learning exploration. After perturbation, the Rayleigh entropy value of Q-learning is 7.4588, and that of the curiosity mechanism is 7.6637, both higher than the input signal's 6.8229. Since a higher Rayleigh entropy value implies lower time-frequency clustering and less information, the curiosity mechanism is superior and can more effectively achieve perturbation.

[0055] (6) Waveform optimization simulation based on transfer learning provided in this embodiment This embodiment uses a model-based transfer learning method, which transfers knowledge learned from the source domain to the target domain with less data, thereby improving the performance of the target domain task. The effectiveness of transfer learning depends mainly on the size of the new dataset and the similarity between the new dataset and the original dataset. Since a reward value matrix Q table is obtained in the Q-learning-curiosity mechanism model, the Q table is frozen and then applied to the actual scenario. This embodiment assumes that the input signal in the actual radar working scenario is as shown in equation (11): (11) The initial Q-table of transfer learning is the Q-table obtained by the Q-learning-curiosity mechanism model, that is, the Q-table obtained by the Q-learning-curiosity mechanism model is frozen, and then the best interference waveform is selected according to the greedy strategy. In transfer learning, the Q value is updated according to formula (12).

[0056] (12) For the next state Possible actions to choose from. This represents the learning rate. This indicates the discount rate.

[0057] The simulation results are shown below, and the final state is as follows. =1226, selected action =1226, and the corresponding parameters of the waveform library are shown in Table 3: Table 3. Waveform library parameters obtained from transfer learning

[0058] The interference results of transfer learning are as follows Figure 9 As shown, it can be seen that after the jamming signal obtained by transfer learning is executed (i.e., after being superimposed with the input signal), the original input signal is basically invisible. In other words, the radar cannot obtain the information it originally wanted, thus achieving the effect of suppression jamming.

[0059] In transfer learning, starting from the second iteration, if the Rayleigh entropy value of the interference signal in the current iteration exceeds 2% of the value in the previous iteration, assuming the Rayleigh entropy value of the interference waveform obtained in the previous exploration was... The Rayleigh entropy value of the signal obtained in this exploration is... ,when If the desired outcome is reached, the corresponding mechanism is triggered. This process continues until the required number of iterations is reached. In this embodiment, the solution with the highest Rayleigh entropy among all curiosity solutions is taken as the output of the curiosity mechanism, i.e., the optimal solution obtained by the curiosity mechanism. Then, the output obtained based on the curiosity mechanism is compared with the original transfer learning to obtain the output result of the transfer learning based on the curiosity mechanism.

[0060] During 100 iterations, the curiosity mechanism was triggered 84 times. The curiosity solution obtained by triggering the curiosity mechanism in transfer learning is as follows: Figure 10 As shown: As shown in the figure above, the signal obtained by executing the curiosity mechanism can suppress the original input signal after interfering with the input signal.

[0061] (II) Generating radar jamming signals based on curiosity-driven Q-learning and LSTM prediction Long Short-Term Memory (LSTM) networks, through their time-series modeling capabilities, can effectively capture the temporal characteristics of radar jamming waveforms, balancing historical patterns with new changes when generating jamming waveforms. After training, LSTMs can generate jamming signals similar to radar waveform characteristics, thereby enhancing jamming effectiveness and improving electronic countermeasures capabilities. An LSTM consists of three gates: a forget gate, an input gate, and an output gate. The forget gate operates as follows: ... Hidden state information at any moment With this moment Data Common input In the function, the output value is between between, The smaller the value output by the function, the more data should be forgotten; the larger the value, the more data should be remembered. The value output by the function is the same as the value at the previous time (i.e. (moment) memory The effect of memory from the previous moment on memory in subsequent moments is limited by multiplication, as shown in equation (13): (13) in It is a measure of the importance of past memories. It is the input at time t. It is the data from the previous hidden layer. It is used to Adjust the matrix to have the same dimensions as the hidden layer at time t. For bias. The working process of the input gate is shown in equations (14) and (15): (14) (15) first Enter information at any time Hidden information from the previous moment After a The function will then the current Enter information at any time Hidden information from the previous moment After passing through a tanh layer, a new candidate value vector is created; then the two values ​​are multiplied to determine the amount of input information retained at the current time step.

[0062] The LSTM-optimized waveform simulation design includes the following steps: The data is divided into an 80% training set and a 20% test set. An LSTM network is constructed with one-dimensional input and output, 300 hidden layers, 500 iterations, and an initial learning rate of 0.005, decaying after 125 epochs. After training, two prediction schemes are used: the first predicts the last 20% of the data step-by-step based on the training set, generating an interference signal y_train_test; the second predicts the entire length step-by-step based on the test set, generating an interference signal y_predict. Finally, the two interference signals are superimposed with a curiosity-based transfer learning interference signal, and the Rayleigh entropy is calculated. The larger value is taken as the final output of the LSTM. Simulation results are as follows: Figure 11 As shown.

[0063] Table 4 Comparison of Rayleigh entropy values ​​in LSTM simulation results

[0064] As shown in Table 4 above, among the two prediction results obtained by the two LSTM schemes, the Rayleigh entropy value of the interference signal obtained by the first scheme is 10.3040, and the Rayleigh entropy value after interfering with the transfer learning input signal using this interference signal is 9.9609. The Rayleigh entropy value of the interference signal obtained by using 100% of the predicted values ​​is 9.4724, and the Rayleigh entropy value after interfering is 9.2759. The final interference signal obtained by using LSTM prediction has the best effect.

[0065] The Rayleigh entropy values ​​of the algorithm results are shown in Tables 5 and 6: Table 5. Rayleigh entropy values ​​of Q-learning results based on curiosity mechanism.

[0066] Table 6. Comparison of Rayleigh Entropy Values ​​in Transfer Learning, Curiosity, and LSTM

[0067] Ultimately, in practical applications of input signal transfer, the output result is as follows: Figure 12 As shown: This algorithm employs a stepwise optimization process. First, it performs transfer learning based on a Q-learning-curiosity mechanism model to obtain an initial interference waveform. Then, a curiosity mechanism is introduced into the transfer learning to explore better solutions, resulting in an optimized interference waveform. This waveform is used as input to an LSTM and further optimized using two prediction schemes. This verifies that the interference waveform obtained through curiosity-based transfer learning is on a gradient trajectory converging towards the optimal solution. Finally, LSTM is used to predict potential better solutions. Experimental results show that the final generated interference signal completely covers the original input signal, and the Rayleigh entropy value increases significantly, indicating a reduction in signal information. This achieves suppressive interference of the input signal, demonstrating the effectiveness of the method in this embodiment.

[0068] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for generating radar jamming signals based on curiosity-driven Q-learning and LSTM prediction, characterized in that, Includes the following steps: S10: Select initial interference waveform parameters from the simulated waveform library based on Q-learning. Rayleigh entropy is used as instantaneous reward and punishment feedback for Q-learning. The action selection is carried out iteratively to update the Q-value table through a greedy strategy to obtain the initial optimal interference waveform parameters. S20: During Q-learning, monitor the change of Rayleigh entropy value of the interference signal in each iteration. When the difference between the Rayleigh entropy value of the current step and the Rayleigh entropy value of the previous step exceeds the threshold, it is judged as an error. Within the error range, explore new actions outside the waveform library and screen the interference signal with the highest Rayleigh entropy to optimize waveform selection performance. S30: Freeze the trained Q-value table and migrate it to the new electromagnetic environment. Using the frozen Q-value table as the initial value, select actions through a greedy strategy and fine-tune parameters according to the Q-value update formula to generate interference waveforms adapted to the new environment. S40, the interference waveforms obtained in steps S20 and S30 are input into the long short-term memory network for prediction optimization. The predicted interference signal is generated by two methods: gradually predicting part of the data in the training set and predicting the full data in one step in the test set. The predicted interference signal is then superimposed with the interference waveforms obtained in step S30, and the signal with the higher Rayleigh entropy is selected as the final optimized interference signal.

2. The radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction as described in claim 1, characterized in that, Step S10 specifically involves: constructing an analog waveform library using linear frequency modulated signals and square waves; using Rayleigh entropy as the instantaneous reward and punishment feedback basis for Q-learning; combining generalized linear frequency modulated transform and time redistribution multi-synchronous compression transform for time-frequency analysis; and using a greedy strategy to select actions to iteratively update the Q-value table to obtain the initial optimal interference waveform parameters.

3. The radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction as described in claim 2, characterized in that, The linear frequency modulated signal is represented as: Square waves are represented as: In the formula, FA is the square wave amplitude, and DR is the duty cycle. Let K be the carrier frequency and K be the linear modulation frequency of the signal. The form of the interference signal output by this waveform library is: .

4. The radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction as described in claim 1, characterized in that, In step S20, it is assumed that the Rayleigh entropy value of the interference waveform detected in the previous step is... ( i>1 The Rayleigh entropy value of the interference signal detected in the current step is... When satisfied At that time, the curiosity mechanism is triggered.

5. The radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction as described in claim 1, characterized in that, In step S30, the Q-value update formula used for fine-tuning the parameters is: In the formula, In order to be in t The state of the surrounding environment at any given moment; In order to be in t At this moment, the environmental state is The action performed at that time; The learning rate represents the next state; Indicates the learning rate; Indicates the discount rate; r This provides instantaneous feedback for each action.

6. The radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction as described in claim 1, characterized in that, In step S40, when the Long Short-Term Memory Network performs prediction optimization, the input data is divided into a training set and a test set in an 8:2 ratio. The training set is used to predict 20% of the data step by step to generate a first prediction interference signal. The test set is used to predict the entire dataset in one step to generate a second prediction interference signal. The Rayleigh entropy of the first prediction interference signal and the second prediction interference signal is calculated, and the signal with the higher Rayleigh entropy is selected as the final interference signal.

7. The radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction as described in claim 1, characterized in that, In step S40, the long short-term memory network includes a forget gate, an input gate, and an output gate. The forget gate controls the retention or forgetting of information in the cell state, the input gate controls the updating of the cell state by the current input information, and the output gate controls the influence of the cell state on the hidden state.

8. The radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction as described in claim 7, characterized in that, In step S40, the working steps of the forget gate are as follows: ... t Hidden state information at time 1 With the current moment t Data Common input The function outputs a value between 0 and 1, and then... The value output by the function is the same as the memory from the previous moment. To limit the influence of memories from a previous moment on memories from subsequent moments by multiplying them, the formula is as follows: In the formula, It is a measure of the importance of past memories. It is the input at time t. It is the data from the previous hidden layer. It is used to Adjust the matrix to have the same dimensions as the hidden layer at time t. For bias.

9. The radar jamming signal generation method based on curiosity-driven Q-learning and LSTM prediction as described in claim 7, characterized in that, The working process of the input gate is as follows: First Enter information at any time Hidden information from the previous moment After a The function will then the current Enter information at any time Hidden information from the previous moment After passing through a tanh layer, a new candidate value vector is created; then the two values ​​are multiplied to determine the amount of input information retained at the current time step.

10. An application of the method as described in any one of claims 1 to 9, characterized in that, It is applied in the fields of military electronic warfare or civilian radar.