Radio frequency modulation fuze intelligent anti-interference decision-making method and system based on reinforcement learning
By introducing a dual Q-learning reinforcement learning algorithm in the fuze system, using two independent Q tables for action selection and reward evaluation, dynamically adjusting the exploration probability, the problem that traditional fuze anti-interference is difficult to make precise decisions in complex environments, and more efficient and accurate anti-interference decisions are achieved.
Patent Information
- Application Number
- CN202510325494.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-20
AI Technical Summary
Traditional fuze anti-interference relies on manual operation, making it difficult to achieve intelligent and precise interference confrontation. Especially in complex and changeable interference environments, it is difficult to accurately estimate future rewards, resulting in frequent fall into local optimal solutions.
A radio frequency modulation fuze intelligent anti-interference decision-making method based on dual Q-learning reinforcement learning algorithm is proposed. By deploying two independent Q tables for action selection and reward evaluation, the exploration probability is dynamically adjusted to dynamically select matching anti-interference measures in complex environments.
It effectively overcomes the overestimation phenomenon that may occur in traditional Q-learning, improves the stability and accuracy of the learning process, improves the reliability and anti-interference performance of the fuze system in complex environments, and the decision-making accuracy rate can be increased by about 20% compared with traditional Q-learning.
Smart Images

Figure CN120185745A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an intelligent anti-jamming decision-making method and system for radio frequency modulation fuzes based on reinforcement learning, belonging to the technical fields of wireless communication and electronic countermeasures. Background Art
[0002] As a key component of a new type of weapon system, a radio frequency modulation fuze relies on radio waves to transmit signals. When the weapon approaches the target, it triggers the explosive device by receiving a specific frequency signal, effectively ensuring the efficient cooperation between the complex battlefield environment and the weapon system. However, the modern battlefield electromagnetic environment presents high complexity and dynamics. With the continuous evolution of intelligent electronic jamming technology and reconnaissance means, the diversity of jamming scenarios and the complexity of anti-jamming methods have increased exponentially. Traditional fuze anti-jamming relies on manual operation, with obvious limitations, and it is difficult to achieve intelligent and precise anti-jamming confrontation. Therefore, it is urgent to study precise anti-jamming strategies for fuzes. With the rise of machine learning technology today, fuze intelligent anti-jamming technology has emerged, and intelligent decision-making is the core element among them. The research on intelligent anti-jamming decision-making methods has become a crucial part of fuze anti-jamming technology. To meet the actual needs of radio frequency modulation fuze intelligent anti-jamming, the present invention introduces a reinforcement learning algorithm and proposes an intelligent anti-jamming decision-making method based on reinforcement learning, which can provide new methods and technical paths for the breakthrough and development of fuze intelligent anti-jamming technology.
[0003] Wang Husheng et al. (H. Wang, B. Chen, Q. Ye, "Design of anti-jamming decision-making for cognitive radar". IET Radar Sonar Navig. vol. 18, no. 3, pp. 514–531, Aug. 2023.) studied the application and implementation of Q-learning, SARSA, and Monte Carlo algorithms in radar anti-jamming decision-making. This literature shows that although the Q-learning algorithm has a fast convergence speed and small reward fluctuations, there are also some deficiencies. For example, the overestimation problem is particularly prominent. During the operation of the Q-learning algorithm, it selects the action corresponding to the next state based on the Q-table and the greedy strategy. Specifically, it tends to overestimate the action that can obtain the maximum reward in the current state, and this phenomenon will be further amplified in a complex and changing anti-jamming confrontation environment. Taking the fuze anti-jamming scenario as an example, the attacker will continuously and flexibly adjust its jamming strategy according to the anti-jamming measures taken by the fuze, and the environmental changes are extremely complex and unpredictable. Affected by its own mechanism and this complex confrontation environment, the Q-learning algorithm is difficult to accurately estimate future rewards, and thus frequently falls into the local optimal solution, ultimately unable to effectively explore the optimal anti-jamming strategy, greatly reducing the anti-jamming performance of the entire system.
[0004] With the continuous development of fuze jamming technology, jammers are more powerful, and the generated fuze jamming patterns are more diverse. It is usually difficult to cope with complex and changeable jamming patterns through artificially designed anti-jamming strategies. Therefore, there is an urgent need for fuzes to have the ability to accurately perceive the environment and stronger intelligent anti-jamming capabilities. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention proposes an intelligent anti-jamming decision-making method for radio frequency modulation fuzes based on the double Q-learning reinforcement learning algorithm. Its core purpose is to enable the fuze system to dynamically and accurately select matching anti-jamming measures when facing a complex electromagnetic environment, so as to ensure the reliable operation of the fuze system in a complex environment. At the same time, this method has the ability to dynamically adjust the exploration probability according to environmental feedback, which can accelerate the model convergence speed and make the system tend to the optimal strategy faster, and can improve its adaptability in a complex and changeable environment.
[0006] The present invention also provides an intelligent anti-jamming decision-making system for radio frequency modulation fuzes based on reinforcement learning.
[0007] Term Explanation:
[0008] 1. Markov decision process (MDP): It is a mathematical framework for describing stochastic decision-making problems. MDP is mainly used to describe decision-making processes with randomness, uncertainty, and multi-step decisions. In MDP, the decision-making problem is modeled as a mathematical framework including state s, action a, reward r, and state transition probability. The state transition in MDP satisfies the Markov property, that is, the future state is only related to the current state and has nothing to do with the past state.
[0009] 2. Reinforcement learning (RL): It is a method for dealing with dynamic decision-making problems. By interacting with the environment through an intelligent agent, different decisions are made to achieve the optimal or specific goal. The value function algorithm is one of the important algorithms in reinforcement learning, and Q-learning is a typical representative of the value function algorithm.
[0010] 3. ε-greedy Strategy: In reinforcement learning, an agent needs to balance exploring unknown actions and exploiting known high-reward actions. The ε-greedy strategy achieves this balance by introducing a parameter ε (epsilon). Specifically, when choosing an action each time, the agent generates a random number between 0 and 1. If this random number is less than ε, the agent randomly selects an action, which is the exploration phase. It gives the agent the opportunity to try new actions and thus discover potentially better strategies. If the random number is greater than or equal to ε, the agent selects the action that it currently believes has the highest estimated value, which belongs to the exploitation phase, that is, using existing experience to obtain the maximum immediate reward. As the training progresses, the value of ε is usually gradually decreased. That is, the agent explores more in the initial stage of learning and relies more on the learned experience in the later stage, hoping to converge to the optimal strategy.
[0011] 4. Q-learning (QL): Q-learning is a model-free reinforcement learning algorithm. It learns the value of taking different actions in a given state by continuously exploring the environment and updating a value called the Q-table. Each entry Q(s,a) in the Q-table represents the expected reward for performing action a in state s. The core of the algorithm is the Q-value update rule, which adjusts the Q-value in the current state-action pair based on the immediate reward and the maximum expected reward of the subsequent state, with the aim of finding the optimal strategy, that is, choosing the action that maximizes the long-term reward in each state.
[0012] The technical solution of the present invention is as follows:
[0013] A radio frequency modulation fuze intelligent anti-jamming decision-making method based on reinforcement learning, comprising:
[0014] Model construction: Construct a state set covering various fuze jamming types and an action set of various anti-jamming measures; construct a reward matrix, and the state transition is carried out randomly to simulate the uncertainty in the real scenario;
[0015] Training and obtaining the best anti-jamming strategy: Use two Q-tables. When updating the Q-value, randomly select one of the Q-tables for action selection and dynamically adjust the exploration probability according to the training process. The other Q-table plays an auxiliary role in the process of updating the Q-value and is used to determine the best action for the next state; by calculating the target Q-value and updating the selected Q-table, while obtaining the reward and updating the decision-making accuracy rate, the two Q-tables alternately perform the above operations and continuously iterate until the Q-value converges, and finally obtain the best anti-jamming strategy for dealing with different interferences.
[0016] Preferably according to the present invention, the state set includes blocking interference, swept-frequency interference, noise frequency modulation interference, dense false target interference, velocity dragging interference, range-velocity dragging interference; the action set includes frequency agility, least mean square algorithm, pseudo-random code modulation, multi-pulse accumulation, phase coding.
[0017] Preferably according to the present invention, a reward matrix is constructed, and corresponding reward values are set for different action combinations under each state.
[0018] Preferably according to the present invention, training and obtaining the optimal anti-interference strategy; including:
[0019] (1) Initialize the Q-table;
[0020] Randomly initialize the Q-table, create two two-dimensional tensors Q1 and Q2 according to the number of states and the number of actions. Q1 and Q2 are two independent Q-tables, used to store the estimated value of each state-action pair, that is, the Q-value; in the Q-table, the rows represent states and the columns represent actions; and initialize Q1 and Q2 with a relatively small positive number. The specific method is: generate a random number in the range of [0, 1), and multiply this random number by 0.01 to change the range to [0, 0.01);
[0021] The Q-table records various information, including: Q-value change, action selection times, correct decision times, total decision times, and decision accuracy rate under each state; Q-value change, that is, record the change of the Q-value of each state-action pair with the training rounds; action selection times, the change of the number of times each action under each state is selected with the progress of training; correct decision times, count the number of times the optimal action is selected under each state; total decision times, record the total number of decisions made under each state; decision accuracy rate under each state, that is, the ratio of the correct decision times to the total decision times;
[0022] (2) Training loop stage;
[0023] In each round of training, the training process includes:
[0024] a. Dynamically adjust the training parameters: According to the current training round number, calculate the current exploration probability ε through the get_epsilon function;
[0025] b. Select the Q-table and actions; including:
[0026] First, randomly determine the Q-table Q selected used to select actions and the Q-table Q other used to assist in calculating the target Q-value, that is, generate a random number of 0-1 through a random generator. If it is less than 0.5, then Q selected is Q1, Q other is Q2, otherwise, then Q selected is Q2, Qother is Q1;
[0027] Next, generate a random number between 0 and 1. If the random number is less than the current exploration probability ε, randomly select an action with uniform probability from all actions. If it is greater than or equal to ε, select the action with the largest Q value in the table for the current state; selected the action with the largest Q value in the table for the current state;
[0028] (3) The agent executes the selected action, and the environment feedbacks an immediate reward to the agent according to the preset reward mechanism. At the same time, the agent enters the next state;
[0029] (4) Convergence judgment and result analysis stage: After the training ends, two Q-tables are finally formed to show the state-action value function learned by the agent; based on the two Q-tables, dynamically match the interference state and the optimal anti-interference action, that is, the optimal anti-interference strategy.
[0030] According to the preference of the present invention, the current exploration probability ε is calculated through the get_epsilon function; including:
[0031] First, multiply the initial exploration probability ε start by the episode power of the decay factor δ to obtain the decayed exploration probability, and then take the maximum value of this decayed exploration probability and the minimum exploration probability ε min as the final exploration probability ε, that is:
[0032] ε = max(ε min , ε start ×δ episode )
[0033] where the get_epsilon function is ε start ×δ episode , ε start is the initial exploration probability, ε min is the minimum exploration probability, and δ is the decay factor of the exploration probability.
[0034] According to the preference of the present invention, in the selection of the Q-table and actions, the probability distribution P(a|s,ε) of action selection is determined by the following formula:
[0035]
[0036] where, is the size of the action set, that is, the number of actions to be selected. Q(s,a) is the state-action value function, representing the expected cumulative discounted reward that the radio fuze system can obtain in the future after executing action a in state s. The algorithm approximates the optimal strategy by continuously updating the Q value; a = argmax a Q(s,a) represents selecting the action a with the largest Q value in the current state s.
[0037] Preferably according to the present invention, the agent (radio fuze system) executes a selected action, and the environment feedbacks an immediate reward to the agent according to a preset reward mechanism, and at the same time the agent enters the next state; including:
[0038] First, through Q other Determine the best action for the next state, and at the same time, based on the immediate reward r, the discount factor γ, and the Q other The Q value corresponding to the best action for the next state in, calculate the target Q value Q target as follows:
[0039] Q target = r + γ * Q other [s′, a′];
[0040] Then, based on the selected Q selected , use the current learning rate to update Q according to the Q-learning update formula selected The Q value Q in the corresponding state-action pair selected (s, a), the update is expressed as follows:
[0041] Q selected (s, a) ← Q selected (s, a) + α · (Q target - Q selected (s, a))
[0042] where α is the learning rate, s and a are the current state and the selected action respectively, s′ is the next state, and a′ is the action selected for the next state.
[0043] A computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the intelligent anti-jamming decision method for radio frequency modulation fuzes based on reinforcement learning.
[0044] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, it implements the steps of the intelligent anti-jamming decision method for radio frequency modulation fuzes based on reinforcement learning.
[0045] The intelligent anti-jamming decision system for radio frequency modulation fuzes based on reinforcement learning includes:
[0046] A model construction module, configured to: construct a state set covering various fuze interference types and an action set of various anti-jamming measures; construct a reward matrix, and the state transition is carried out randomly to simulate the uncertainty in the real scenario;
[0047] The training and optimal anti-interference strategy acquisition module is configured to: use two Q-tables. When updating the Q-value, randomly select one of the Q-tables for action selection and dynamically adjust the exploration probability according to the training process. The other Q-table plays an auxiliary role in the process of updating the Q-value and is used to determine the optimal action for the next state; by calculating the target Q-value and updating the selected Q-table, while obtaining the reward and updating the decision-making accuracy rate, the two Q-tables alternately perform the above operations and continuously iterate until the Q-value converges, and finally obtain the optimal anti-interference strategy for dealing with different interferences.
[0048] The beneficial effects of the present invention are as follows:
[0049] 1. The double Q-learning algorithm effectively overcomes the overestimation phenomenon that may occur in traditional Q-learning by deploying two independent Q-tables for action selection and reward evaluation respectively. In the standard Q-learning, since the action selection and value evaluation share one Q-table, it is easy to cause the overestimation of the value of some state-action pairs, thus affecting the learning efficiency and algorithm performance. The double Q-learning algorithm improves the stability and accuracy of the learning process by separating these two functions.
[0050] 2. By adopting the ε-greedy strategy to dynamically adjust the parameters, the agent can fully explore the environment at the initial stage of training and make more effective decisions based on the accumulated knowledge in the later stage of training. At the same time, this strategy also ensures that the agent can still maintain good adaptability and flexibility when facing unknown or dynamically changing environments. This ability is crucial in practical applications, especially when dealing with diverse time-varying interference signals, the agent needs to have the ability to quickly adapt to and respond to new situations.
[0051] 3. The simulation results show that the decision-making accuracy rate of this patent can be increased by about 20% at most compared with traditional Q-learning, that is, the fuze system will more accurately and effectively select the optimal anti-interference strategy when identifying and dealing with various interference signals. The above advantages benefit from the advanced mechanism of the double Q-learning algorithm, as well as the constructed reward matrix and parameter optimization strategy, which jointly ensure that the fuze system quickly decides the optimal strategy during the training process. Description of the Drawings
[0052] Figure 1 is a schematic flow chart of the intelligent anti-interference decision-making method of the radio frequency modulation fuze based on reinforcement learning of the present invention.
[0053] Figure 2 is a schematic diagram of the Q-value change of each anti-interference measure when the double Q-learning of the present invention resists the swept-frequency interference.
[0054] Figure 3It is a schematic diagram of the Q-value change of each anti-jamming measure when traditional Q-learning confronts swept-frequency interference.
[0055] Figure 4 It is a schematic diagram of the change in the number of selection decisions of each anti-jamming measure when the dual Q-learning of the present invention confronts swept-frequency interference.
[0056] Figure 5 It is a schematic diagram of the change in the number of selection decisions of each anti-jamming measure when traditional Q-learning confronts swept-frequency interference.
[0057] Figure 6 It is a histogram of the correct rate of the dual Q-learning anti-jamming measures.
[0058] Figure 7 It is a histogram of the correct rate of the traditional Q-learning anti-jamming measures. Specific implementation mode
[0059] The present invention will be further defined below in conjunction with the accompanying drawings of the specification and implementation cases, but not limited thereto.
[0060] Example 1
[0061] The intelligent anti-jamming decision-making method for radio frequency modulation fuzes based on reinforcement learning includes:
[0062] Model construction: Construct a state set covering various types of fuze interference and an action set of various anti-jamming measures; construct a reward matrix, and the state transition is carried out randomly to simulate the uncertainty in the real scenario;
[0063] Training and obtaining the best anti-jamming strategy: Use two Q-tables. When updating the Q-value, randomly select one of the Q-tables for action selection, and dynamically adjust the exploration probability according to the training process. The other Q-table plays an auxiliary role in the process of updating the Q-value and is used to determine the best action for the next state; by calculating the target Q-value and updating the selected Q-table, at the same time obtaining the reward and updating the decision-making correct rate, the two Q-tables alternately execute the above operations, continuously iterate until the Q-value converges, and finally obtain the best anti-jamming strategy for dealing with different interferences. Achieve the effect of giving the optimal decision for various interferences.
[0064] Example 2
[0065] The intelligent anti-jamming decision-making method for radio frequency modulation fuzes based on reinforcement learning according to Example 1 is characterized in that:
[0066] The state set includes blocking interference, frequency sweep interference, noise frequency modulation interference, dense false target interference, velocity dragging interference, and range-velocity dragging interference; the action set includes frequency agility, least mean square (LMS) algorithm, pseudo-random code modulation, multi-pulse accumulation, and phase coding. These constitute part of the environmental model that the agent needs to learn.
[0067] Define the state set of interference signals and the action set of anti-interference measures to clarify all possible situations in the system. For example, various fuze interference types, such as blocking interference and frequency sweep interference, constitute the state set, and corresponding actions that can be taken, such as frequency agility and least mean square (LMS) algorithm, form the action set, thereby determining the scope of the decision-making environment where the agent is located.
[0068] In addition, to implement the double Q-learning algorithm, two Q-tables Q1 and Q2 are randomly initialized with small positive values to avoid all initial values being zero, which may lead to decision-making bias in the initial exploration stage. At the same time, key parameters are set, including the learning rate, discount factor, and parameters related to the ε-greedy strategy, such as the initial exploration probability, minimum exploration probability, and decay factor. The selection of these parameters is crucial for the learning efficiency and final performance of the agent.
[0069] Construct a reward matrix and set corresponding reward values for different action combinations in each state. Combining prior knowledge analysis, although there are multiple methods to complete the task that needs to be executed, there is only one optimal method. Therefore, only the optimal action in each state is given a relatively high positive reward, and the remaining actions are given lower rewards. In this way, the agent is guided to learn towards the desired optimal strategy. The specific reward assignment is shown in Table 1.
[0070] Table 1
[0071]
[0072] Training and obtaining the optimal anti-interference strategy; including:
[0073] (1) Initialize the Q-table;
[0074] Randomly initialize the Q-table. Create two two-dimensional tensors Q1 and Q2 according to the number of states and the number of actions. Q1 and Q2 are two independent Q-tables used to store the estimated values, i.e., Q-values, of each state-action pair; their shapes are determined by the number of states and the number of actions. In the Q-table, the rows represent states and the columns represent actions. For example, when there are 3 states and 2 actions, both Q1 and Q2 are two-dimensional tensors of (3,2). And Q1 and Q2 are initialized with small positive numbers. The specific method is: generate a random number in the range of [0,1), and multiply this random number by 0.01 to change the range to [0,0.01); avoid the initial Q-value being too large to affect learning and let the algorithm start learning from a relatively fair starting point.
[0075] Two Q-tables are used to store the estimated values of each state-action pair, namely Q-values, and continuously update and cooperate with each other during the subsequent learning process;
[0076] The Q-table records multiple aspects of information, including: Q-value changes, action selection times, correct decision-making times, total decision-making times, and decision-making accuracy rates for each state; Q-value changes, that is, record the changes in the Q-values of each state-action pair with the training rounds; action selection times, the changes in the number of times each action is selected for each state during training; correct decision-making times, count the number of times the optimal action is selected for each state; total decision-making times, record the total number of decisions made for each state; decision-making accuracy rates for each state, that is, the ratio of the correct decision-making times to the total decision-making times; these records help to comprehensively analyze the learning process and final effect of the double Q-learning algorithm.
[0077] It is convenient for subsequent analysis and visual display of the learning process.
[0078] (2) Training loop stage (for each training episode);
[0079] In each round of training, the training process includes:
[0080] a. Dynamically adjust training parameters: According to the current training round (episode), calculate the current exploration probability ε through the get_epsilon function; including:
[0081] First multiply the initial exploration probability ε start (with a value of 1.0) by the decay factor δ (with a value of 0.995) to the power of episode to obtain the decayed exploration probability, and then take the maximum value of this decayed exploration probability and the minimum exploration probability ε min (with a value of 0.01) as the final exploration probability ε, so that the algorithm has a high probability of exploration in the initial stage of training. As the number of training rounds increases, the exploration probability gradually decreases, but will not be lower than the minimum exploration probability, to balance exploration and exploitation. That is:
[0082] ε = max(ε min , ε start ×δ episode );
[0083] Among them, the get_epsilon function is ε start ×δ episode , ε start is the initial exploration probability, ε min$\epsilon$ is the minimum exploration probability, and $\delta$ is the decay factor of the exploration probability. This decay factor is affected by time factors. Through such dynamically changing parameter settings, the learning process can exhibit stage characteristics: in the early stage, the agent will focus on exploring new actions to discover as many possibilities as possible; as training progresses, in the later stage, the agent will be more inclined to utilize the previously learned better strategies to optimize learning efficiency and effects. After completing the dynamic adjustment of training parameters, it is then necessary to ensure that each state has the opportunity to be selected in multiple different training rounds, so that the model can effectively simulate the real process of the agent continuously learning and accumulating experience in various environmental states.
[0084] b. Select the Q-table and actions; including:
[0085] In the double Q-learning algorithm, first, randomly determine the Q-table $Q$ selected used to select actions and the Q-table $Q'$ other used to assist in calculating the target Q value, that is, generate a random number between 0 and 1 through a random generator. If it is less than 0.5, then $Q$ selected is $Q_1$, $Q$ other is $Q_2$, otherwise, $Q$ selected is $Q_2$, $Q$ other is $Q_1$;
[0086] Then, select an action based on the $\epsilon$-greedy policy ($\epsilon$ update method in step a). Generate another random number between 0 and 1. If the random number is less than the current exploration probability $\epsilon$, randomly select an action with uniform probability from all actions. If it is greater than or equal to $\epsilon$, select the action with the largest Q value in the current state in the $Q$ selected table; this balances exploring new actions and exploiting known optimal actions;
[0087] Determine the probability distribution $P(a|s,\epsilon)$ of action selection through the following formula:
[0088]
[0089] where, $|A|$ is the size of the action set, that is, the number of available actions, $Q(s,a)$ is the state-action value function, representing the expected cumulative discounted reward that the radio fuze system can obtain in the future after performing action $a$ in state $s$. The algorithm approaches the optimal policy by continuously updating the Q value; $a = \arg\max$ a $Q(s,a)$ represents selecting the action $a$ with the largest Q value in the current state $s$. The $\epsilon$-greedy policy ensures that the agent neither locks in certain actions too early nor completely ignores the learned knowledge, but gradually converges to the optimal policy in a balanced manner.
[0090] (3) The agent (radio fuze system) executes the selected action, and the environment feedbacks an immediate reward to the agent according to the preset reward mechanism. Meanwhile, the agent enters the next state, including:
[0091] First, determine the best action for the next state through Q other At the same time, based on the immediate reward r, the discount factor γ, and the Q value Q corresponding to the best action for the next state in Q other Calculate the target Q value Q target as follows:
[0092] Q target = r + γ * Q other [s′, a′];
[0093] Then, based on the selected Q selected , update Q using the current learning rate according to the Q-learning update formula selected The Q value Q in the corresponding state-action pair in Q selected (s, a), and the update is shown as follows:
[0094] Q selected (s, a) ← Q selected (s, a) + α · (Q target - Q selected (s, a))
[0095] where α is the learning rate, s and a are the current state and the selected action respectively, s′ is the next state, and a′ is the action selected for the next state.
[0096] In the above way, the action selection and value evaluation can be effectively separated, thereby reducing the overestimation problem caused by a single Q-table. The double Q-learning algorithm reduces the overestimation problem by introducing two independent Q-tables, making the learning process more stable and improving the quality of the final decision.
[0097] (4) Convergence judgment and result analysis stage: After the training is completed, two Q-tables are finally formed to show the state-action value function learned by the agent; based on the two Q-tables, the interference state and the optimal anti-interference action are dynamically matched, that is, the optimal anti-interference strategy.
[0098] In addition, a series of other charts are also generated, including the Q value change chart, the action selection times change chart, and the decision accuracy histogram. Through this histogram, the accuracy of the learned strategies under different interference states can be intuitively compared;
[0099] By introducing two independent Q-tables, the above algorithm effectively separates the operation of action selection and value evaluation. In the traditional single Q-table learning process, overestimation problems often occur due to its own limitations. The proposed double Q-learning algorithm cleverly avoids this problem. Since the two Q-tables cooperate with each other and are independent of each other, in the learning process, one Q-table focuses on action selection, and the other focuses on value evaluation and auxiliary update. This division of labor and cooperation mode greatly improves the stability of the entire learning process and is no longer easily trapped in the decision-making deviation caused by overestimation. Based on this, the algorithm trains the Q-table to learn the Q-values of each action under different interference states. Finally, in a complex electromagnetic environment, the radio fuze system can dynamically match the interference state with the optimal anti-interference action according to the Q-table: when encountering "blocking interference" and "frequency sweeping interference", select "frequency agility"; in the face of "noise frequency modulation interference", adopt the "least mean square algorithm"; for "dense false target interference", choose "multiple pulse accumulation"; when dealing with "velocity towing interference" and "range-velocity towing interference", enable "phase coding". Through this precise matching, the efficiency and accuracy of the fuze system's decision-making are significantly improved.
[0100] The simulation results show that the double Q-learning algorithm designed in the present invention effectively reduces the overestimation problem that may exist in single Q-learning by introducing two independent Q-tables to select actions and evaluate rewards, making the learning process more stable. All algorithms of the present invention have been simulated in PyCharm, and the specific experimental results are analyzed as follows:
[0101] Figure 1 It is the flow chart of the double Q-learning algorithm. This figure elaborates in detail the operation steps of the double Q-Learning algorithm. In the initial stage, the system initializes two Q-tables (Q1 and Q2). Then, the system enters the training loop and randomly selects a Q-table (named Q selected ) to determine the action in the current state. This decision is based on the Q-values in the Q selected table and may combine with the dynamic ε-greedy strategy for random exploration. The dynamic ε-greedy strategy enables the system to select non-optimal actions with a certain probability to explore new state-action combinations. At the same time, as the training progresses, the ε value will gradually decrease to balance the exploration and exploitation capabilities of the algorithm. After the action is selected, the system executes the action and collects the immediate reward of the current state. Subsequently, the system transfers to the next state. Before reaching the preset number of training rounds, the system will use the other Q-table (Q other ) and the reward to calculate the target Q-value, and determine the optimal action of the next state based on the Q other table. Then, the system updates the selected Q-table (Q selected) At the same time, follow the update rules of Q-learning and continue the decision-making process for the next round. The system will continuously execute these steps until the number of training rounds reaches the preset total number of rounds. In each round, the system gradually improves the strategy by randomly selecting the Q-table and updating the Q-values, and finally masters the optimal action selection strategy. When the number of training rounds is completed, the system reaches the "end" node and the algorithm terminates accordingly. Through this process, the Double Q-Learning algorithm can continuously optimize the Q-values in multiple iterations, and then learn the best strategy, effectively coordinating the balance between exploration and exploitation.
[0102] Figure 2 It shows the excellent performance of the Double Q-learning algorithm of the present invention in countering swept-frequency interference, where the variation trends of the Q-values of various anti-interference measures (frequency agility, least mean square algorithm, pseudo-random code modulation, multi-pulse accumulation, and phase coding) with the number of countering times are clearly depicted. As can be seen from the figure, the Q-value of the frequency agility strategy (blue line) rises rapidly in the initial stage and finally stabilizes at a relatively high level, indicating the effectiveness and stability of this strategy in dealing with swept-frequency interference. In contrast, although the Q-values of other anti-interference measures such as the least mean square (LMS) algorithm (orange line), pseudo-random code modulation (green line), multi-pulse accumulation (red line), and phase coding (purple line) also increase to some extent in the initial stage, the increase amplitude is small, and they fluctuate greatly in the subsequent countering process and finally fail to reach the Q-value level of the frequency agility strategy. This phenomenon shows that the agent has tried various anti-interference measures during the training process and, through learning the reward function matrix, has gradually recognized the advantages of the frequency agility strategy in countering swept-frequency interference.
[0103] Figure 3Although the traditional Q-learning algorithm presented shows an upward trend in Q-values for various anti-jamming measures, the difference in its improvement speed and the final Q-value indicates the limitations of this algorithm in selecting the optimal strategy. As can be seen from the figure, although the Q-value of the frequency agility strategy finally converges to a relatively high level, the Q-values of other anti-jamming measures are also relatively high, which reflects the bias in the traditional Q-learning algorithm when evaluating the value of strategies. It relies on a single Q-table for action selection and value evaluation, and is prone to falling into the dilemma of overestimation, thus affecting the convergence speed of the algorithm and possibly leading to inaccurate final strategies. In addition, it is difficult for the traditional Q-learning algorithm to achieve a balance between exploration (i.e., trying new actions to discover potentially high-reward strategies) and exploitation (i.e., selecting actions with better past rewards based on existing experience), and it is necessary to carefully adjust the exploration rate to avoid prematurely converging to suboptimal strategies or reducing efficiency due to excessive exploration. The dual Q-learning algorithm of the present invention effectively solves these problems by separating action selection and value evaluation, enabling the algorithm to identify the optimal strategy faster and demonstrating stronger adaptability and more efficient decision-making ability in a complex and changing electromagnetic environment. Figure 2 The results shown not only highlight the significant advantages of the present invention in dynamically selecting anti-jamming measures but also confirm its potential in practical applications.
[0104] Figure 4 and Figure 5 show the specific selection of various anti-jamming measures against swept-frequency jamming. From Figure 5 it can be seen that as the number of confrontations increases, the number of selection decisions for each anti-jamming measure gradually increases, but the increase is relatively gentle. This indicates that the traditional Q-learning algorithm has a slow learning speed when facing a complex electromagnetic environment and is difficult to quickly adapt to new jamming patterns. However Figure 4 the performance of the dual Q-learning algorithm is quite different. As the number of confrontations increases, the number of selection decisions for each anti-jamming measure shows an obvious upward trend. During the 336 confrontations, the frequency agility was selected 305 times, while the traditional Q-learning algorithm only made 239 frequency agility strategy selections. This shows that the dual Q-learning algorithm can learn and adapt to new jamming patterns faster, especially when facing a complex electromagnetic environment, its decision-making efficiency is higher. In addition, by comparing the two figures, it can also be found that the dual Q-learning algorithm switches more quickly and frequently between different anti-jamming measures, which means that it can more quickly and flexibly select appropriate coping strategies when facing different types of jamming, while the flexibility of the traditional Q-learning algorithm is significantly insufficient.
[0105] Figure 6 and Figure 7The comparison between the dual Q-learning algorithm proposed by the present invention and the traditional Q-learning algorithm in terms of decision-making accuracy is respectively shown. The chart data indicates that in terms of the strategy selection for coping with various interference signals, the dual Q-learning algorithm has obvious advantages over the traditional Q-learning algorithm. Under 6 different interference types, the accuracy rate of the dual Q-learning algorithm is generally higher than that of the traditional Q-learning algorithm, and the maximum increase reaches 20%. Especially when coping with "swept-frequency interference" and "blocking interference", the accuracy rates of the dual Q-learning algorithm are as high as 91% and 93% respectively, while the accuracy rates of the traditional Q-learning algorithm are only 71% and 74% respectively. This comparison result fully confirms that the dual Q-learning algorithm not only surpasses the traditional Q-learning algorithm in overall performance, but also is more prominent in terms of stability and adaptability when facing specific types of interference.
[0106] Example 3
[0107] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the intelligent anti-interference decision-making method for radio frequency modulation fuzes based on reinforcement learning described in Example 1 or 2 are implemented.
[0108] Example 4
[0109] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the intelligent anti-interference decision-making method for radio frequency modulation fuzes based on reinforcement learning described in Example 1 or 2 are implemented.
[0110] Example 5
[0111] An intelligent anti-interference decision-making system for radio frequency modulation fuzes based on reinforcement learning includes:
[0112] A model construction module is configured to: construct a state set covering various fuse interference types and an action set of various anti-interference measures; construct a reward matrix, and the state transition is carried out randomly to simulate the uncertainty in the real scenario;
[0113] A training and optimal anti-interference strategy acquisition module is configured to: use two Q-tables. When updating the Q value, randomly select one of the Q-tables for action selection, and dynamically adjust the exploration probability according to the training process. The other Q-table plays an auxiliary role in the process of updating the Q value and is used to determine the optimal action for the next state; by calculating the target Q value and updating the selected Q-table, and at the same time obtaining the reward and updating the decision-making accuracy rate, the two Q-tables alternately execute the above operations, continuously iterate until the Q value converges, and finally obtain the optimal anti-interference strategy for coping with different interferences.
Claims
1. An intelligent anti-interference decision-making method for radio frequency modulation fuse based on reinforcement learning, characterized in that: include: Model construction: construct a state set covering various types of fuze interference and an action set of various anti-interference measures; The reward matrix is constructed and the state transition is performed in a random manner to simulate the uncertainty in real scenarios; Training and acquisition of the best anti-interference strategy: Use two Q-tables. When updating the Q value, randomly select one of the Q-tables for action selection, and dynamically adjust the exploration probability according to the training process. The other Q-table plays an auxiliary role in the process of updating the Q value and is used to determine the best action for the next state. By calculating the target Q value and updating the selected Q table, rewards are obtained and the decision accuracy is updated. The two Q-tables perform the above operations alternately and iterate continuously until the Q value converges, and finally the best anti-interference strategy for dealing with different interferences is obtained.
2. The intelligent anti-interference decision method for radio frequency modulation fuse based on reinforcement learning according to claim 1 is characterized in that: The state set includes blocking interference, sweeping frequency interference, noise frequency modulation interference, dense false target interference, speed dragging interference, and distance speed dragging interference; the action set includes frequency agility, least mean square algorithm, pseudo-random code modulation, multi-pulse accumulation, and phase coding.
3. The intelligent anti-interference decision method for radio frequency modulation fuse based on reinforcement learning according to claim 1 is characterized in that: Construct a reward matrix and set corresponding reward values for different action combinations in each state.
4. The intelligent anti-interference decision method for radio frequency modulation fuse based on reinforcement learning according to claim 1 is characterized in that: Training and acquisition of the best anti-interference strategy; including: (1) Initialize the Q table; Randomly initialize the Q table, create two two-dimensional tensors Q1 and Q2 according to the number of states and actions, Q1 and Q2 are two independent Q tables, used to store the estimated value of each state-action pair, i.e., the Q value; in the Q table, rows represent states and columns represent actions; and initialize Q1 and Q2 with small positive numbers, specifically: generate a random number in the range of [0,1), and multiply the random number by 0.01 to change the range to [0,0.01); The Q table records various information, including: Q value changes, action selection times, correct decision times, total decision times, and decision accuracy rate in each state; Q value changes, that is, recording the changes of Q value of each state-action pair with the training rounds; action selection times, the number of times each action is selected in each state changes with the training; correct decision times, counting the number of times the optimal action is selected in each state; total decision times, recording the total number of decisions made in each state; decision accuracy rate in each state, that is, the ratio of correct decision times to total decision times; (2) Training cycle phase; In each round of training, the training process includes: a. Dynamically adjust training parameters: According to the current number of training rounds, calculate the current exploration probability ε through the get_epsilon function; b. Select Q table and action; including: First, the Q table Qselected used to select the action and the Q table Qother used to assist in calculating the target Q value are randomly determined. That is, a random number between 0 and 1 is generated by a random generator. If it is less than 0.5, Qselected is Q1 and Qother is Q2. Otherwise, Qselected is Q2 and Qother is Q1. Next, generate a random number between 0 and 1. If the random number is less than the current exploration probability ε, randomly select an action from all actions with uniform probability. If the random number is greater than or equal to ε, select the action with the largest Q value in the current state in the Qselected table. (3) The agent performs the selected action, and the environment gives the agent an immediate reward based on the preset reward mechanism, and the agent enters the next state; (4) Convergence judgment and result analysis phase: After training, two Q-tables are finally formed to display the state-action value function learned by the intelligent agent; based on the two Q-tables, the interference state and the optimal anti-interference action are dynamically matched, that is, the optimal anti-interference strategy.
5. The intelligent anti-interference decision method for radio frequency modulation fuse based on reinforcement learning according to claim 4 is characterized in that: The current exploration probability ε is calculated through the get_epsilon function; including: First, multiply the initial exploration probability εstart by the episode power of the decay factor δ to get the decayed exploration probability, and then take the maximum value of the decayed exploration probability and the minimum exploration probability εmin as the final exploration probability ε, that is: ε=max(εmin,εstart×δ episode ) Among them, the get_epsilon function is εstart×δ episode , εstart is the initial exploration probability, εmin is the minimum exploration probability, and δ is the attenuation factor of the exploration probability.
6. The intelligent anti-interference decision method for radio frequency modulation fuse based on reinforcement learning according to claim 4 is characterized in that: In the Q table and action selection, the probability distribution of action selection P(a|s,ε) is determined by the following formula: in, is the size of the action set, that is, the number of actions to choose from. Q(s,a) is the state-action value function, which represents the expected cumulative discounted reward that the radio fuze system can obtain in the future after performing action a in state s. The algorithm approaches the optimal strategy by continuously updating the Q value; a=argmaxa Q(s,a) means selecting action a with the largest Q value in the current state s.
7. The intelligent anti-interference decision method for radio frequency modulation fuse based on reinforcement learning according to claim 4 is characterized in that: The agent performs the selected action, and the environment gives the agent an immediate reward based on the preset reward mechanism, and the agent enters the next state; including: First, the best action for the next state is determined by Qother. At the same time, based on the immediate reward r, the discount factor γ, and the Q value of the best action for the next state in Qother, the target Q value Qtarget is calculated as follows: Qtarget=r+γ*Qother[s′,a′]; Then, based on the selected Qselected, the current learning rate is used to update the Q value Qselected(s,a) in the corresponding state-action pair in Qselected according to the Q-learning update formula. The update is expressed as follows: Qselected(s,a)←Qselected(s,a)+α·(Qtarget-Qselected(s,a)) Where α is the learning rate, s and a are the current state and the selected action respectively, s′ is the next state, and a′ is the action selected for the next state.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the intelligent anti-interference decision-making method for radio frequency modulation fuse based on reinforcement learning as described in any one of claims 1-7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the steps of the intelligent anti-interference decision-making method for radio frequency modulation fuse based on reinforcement learning as described in any one of claims 1-7 are implemented.
10. The intelligent anti-interference decision-making system for radio frequency modulation fuse based on reinforcement learning is characterized by: include: The model construction module is configured to: construct a state set covering multiple fuze jamming types and an action set covering multiple anti-jamming measures; The reward matrix is constructed and the state transition is performed in a random manner to simulate the uncertainty in real scenarios; The training and optimal anti-interference strategy acquisition module is configured as follows: using two Q-tables, when updating the Q value, randomly selecting one of the Q-tables for action selection, and dynamically adjusting the exploration probability according to the training process, the other Q-table plays an auxiliary role in the process of updating the Q value, and is used to determine the best action for the next state; by calculating the target Q value and updating the selected Q-table, while obtaining rewards and updating the decision accuracy, the two Q-tables perform the above operations alternately, and continue to iterate until the Q value converges, and finally obtain the optimal anti-interference strategy for dealing with different interferences.