Intelligent jamming method based on cooperative Q-learning
By adopting an intelligent interference method based on cooperative Q-learning, coordinated decision-making of multi-agent interference machines is achieved, which solves the problem of insufficient interference effect in multi-agent adversarial scenarios, improves the interference success rate and reduces user throughput loss.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-21
- Publication Date
- 2026-04-07
AI Technical Summary
Existing intelligent jamming technologies are insufficient to cope with multi-agent adversarial scenarios. Single intelligent jamming devices cannot guarantee their own concealment and reliable jamming effect in complex electromagnetic spectrum spaces. The lack of coordination among multiple intelligent jamming devices leads to serious frequency conflicts and greatly reduces the jamming effect.
A cooperative Q-learning-based intelligent jamming method is adopted. Through the cooperation of M intelligent jammers and N pairs of communication users, independent and joint Q-value tables are established. An ε-greedy strategy is used to select joint actions, evaluate the jamming effect, and update the Q-value tables, thereby achieving coordination of decision-making within the jammers.
In multi-agent adversarial scenarios, the jamming effect is improved, and coordinated decision-making within the jammer is achieved. No prior information about the user or channel is required. By interacting with the spectrum environment to optimize strategies online, the jamming success rate is improved and the user throughput loss is reduced.
Smart Images

Figure CN115567148B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless communication technology, specifically a smart interference method based on cooperative Q-learning. Background Technology
[0002] The electromagnetic spectrum is a powerful support for forming joint operational capabilities of network information systems. Based on the development trend and practical needs of communication jamming and countermeasures, conducting research on technologies that can effectively disrupt and damage enemy communications is crucial, and research in the field of communication jamming is becoming increasingly urgent. However, traditional communication jamming methods, such as fixed jamming, frequency sweeping jamming, and comb jamming, have fixed jamming patterns and are difficult to effectively counter dynamic anti-jamming measures. Therefore, in recent years, researchers have continuously proposed intelligent jamming technologies based on machine learning. Empowered by artificial intelligence algorithms, jammers can learn and uncover patterns of user communication changes, thereby adopting efficient and reliable jamming methods. Existing research has applied reinforcement learning methods to the jamming field, proposing an "online perception, virtual decision-making" jamming decision-making method based on reinforcement learning, enabling jammers to effectively learn and jam without prior information from communication users (S. Zhang, H. Tian, X. Chen, et al., "Design and implementation of reinforcement learning-based intelligent jamming system," IET Communications, vol.14, no.18, pp.3231-3238, Nov.2020). Similarly, existing literature has also applied deep reinforcement learning methods to UAV anti-jamming. Jamming UAVs intelligently interfere by observing the trajectories of communication UAVs, while communication UAVs have also designed deep reinforcement learning algorithms to evade attacks from jamming UAVs (N. Gao, Z. Qin, X. Jing, Q. Ni, and S. Jin, “Anti-intelligent UAV jamming strategy via deep Q-networks,” IEEE Transactions on Communications, vol. 68, no. 1, pp. 569-581, 2020.). Furthermore, existing literature employs deep learning-based jammers to predict channel transmission quality, achieving precise jamming, and uses generative adversarial networks to reduce training time with limited samples (T. Erpek, YESagduyu and Y. Shi, “Deep learning for launching and mitigating wireless jamming attacks,” IEEE Transactions on Cognitive Communications and Networking, vol. 5, no. 1, pp. 2-14, Mar. 2019.).However, the above researchers have considered adversarial scenarios based on a single jammer and a pair of communication users. The jammer's intelligent decision-making ability is limited and the communication adversary it faces is not strong. In scenarios where multiple communication users communicate simultaneously, a single intelligent jammer is difficult to cope with multi-agent adversarial environments.
[0003] On the other hand, as reinforcement learning has achieved remarkable results in multiple application areas, and considering that multiple decision-making individuals (agents) usually exist simultaneously in real-world scenarios, some researchers have gradually extended their focus from the single-agent domain to multi-agent domains, namely multi-agent reinforcement learning. Currently, a small number of studies have investigated multi-agent anti-jamming scenarios. Existing literature has considered the coordination among communication users and proposed a collaborative multi-agent anti-jamming algorithm based on RL to obtain the optimal anti-jamming strategy (F. Yao and L. Jia, “A collaborative multi-agent reinforcement learning anti-jamming algorithm in wireless networks,” IEEE Wireless Communications Letters, vol. 8, no. 4, pp. 1024–1027, 2019.). In addition, some literature has proposed a model-free multi-agent reinforcement learning algorithm, which improves Nash Q learning by using the idea of mean-field game theory. It treats all nearby agents as a whole and only cares about the actions of the whole, thus greatly reducing complexity (Yang Y, Luo R, Li M, et al. Mean Field Multi-Agent Reinforcement Learning[C]. The 35th International Conference on Machine Learning, 2018.). Currently, research on cooperative jamming mainly focuses on cooperative deception jamming against radar detection, or friendly jamming to ensure one's own secure communication when facing enemy eavesdropping. Research on multi-domain cooperative jamming that actively disrupts enemy communication is still relatively scarce. Therefore, it is necessary to study jamming strategies applicable to multi-agent adversarial scenarios.
[0004] In summary, existing research on intelligent jamming is insufficient to directly address multi-agent adversarial scenarios, exhibiting the following problems: 1) Single-agent jamming is inadequate for multi-agent adversarial environments. In complex electromagnetic spectrum spaces, enemy communication devices are numerous, their intelligent anti-jamming capabilities are increasingly sophisticated, and communication modes and patterns are dynamically changing, resulting in high spectrum occupancy. Single-agent jamming devices struggle to maintain both their stealth and reliable jamming effectiveness in multi-agent communication environments; 2) Severe frequency conflicts exist within multi-agent jamming systems. In multi-agent communication environments, the goal of jamming devices is to suppress enemy communication devices on the spectrum. Lack of coordination between jammers leads to significant frequency conflicts, a large proportion of ineffective jamming, and a substantial reduction in jamming effectiveness. Therefore, simply superimposing single-agent jamming devices cannot be directly applied to multi-agent adversarial scenarios. Summary of the Invention
[0005] The purpose of this invention is to provide an intelligent interference method based on cooperative Q-learning, which can effectively improve the interference effect in multi-agent adversarial scenarios.
[0006] The technical solution for achieving the objective of this invention is: an intelligent interference method based on cooperative Q-learning, characterized by comprising the following steps:
[0007] Step 1: Consider an interference scenario consisting of M intelligent jammers and N pairs of communication users. In this scenario, the communication users cooperate to determine the communication channel. The users communicate using either a fixed-sequence frequency hopping method or a random-sequence frequency hopping method. The intelligent jammers possess sensing and learning capabilities, and can perceive the real-time environmental spectrum status s. t ;
[0008] Step 2: Each intelligent jammer establishes and maintains two Q-value tables: an independent Q-value table and a joint Q-value table. The jammer takes the sensed user spectrum state as the state input and selects the joint action a = {a1, ..., a2} according to the ε-greedy strategy. M};
[0009] Step 3: Execute the joint action, evaluate the effect of releasing the interference based on the joint action, and obtain the reward value r for each jammer under the current joint action. m (s t ,a m ), and the total gain R of the entire set of disturbances. t (s,a), and sense the current spectrum state s t +1;
[0010] Step 4: Update the independent Q-value table and the combined Q-value table based on the reward values received;
[0011] Step 5: Repeat steps 1 through 4 until the specified number of iterations.
[0012] The present invention adopts the above technical solution and has the following advantages compared with the prior art:
[0013] 1. Focusing on the cutting-edge application background of multi-agent cooperative interference, this study investigates the joint decision-making method of multi-agent interference in multi-agent adversarial scenarios, which can realize the coordination of internal decisions of multi-agent interference machines and effectively improve the interference effect in multi-agent adversarial scenarios.
[0014] 2. Without the need for prior user and channel information, the jammer can continuously optimize its strategy online by interacting with the spectrum environment and learning. Attached Figure Description
[0015] Figure 1 This is a schematic diagram illustrating the adversarial mechanism of the intelligent interference method based on cooperative Q-learning of this invention.
[0016] Figure 2 This is a framework diagram of the intelligent interference method based on cooperative Q-learning of this invention.
[0017] Figure 3 This is a schematic diagram illustrating the interference success probability performance of the method and comparison algorithm proposed in Embodiment 1 of the present invention.
[0018] Figure 4 This is a schematic diagram illustrating the user-normalized throughput performance of the method and comparison algorithm proposed in Embodiment 1 of the present invention.
[0019] Figure 5 This is a schematic diagram illustrating the interference success probability performance of the method and comparison algorithm proposed in Embodiment 2 of the present invention.
[0020] Figure 6 This is a schematic diagram illustrating the user-normalized throughput performance of the method and comparison algorithm proposed in Embodiment 2 of the present invention. Detailed Implementation
[0021] This invention proposes an intelligent jamming method based on cooperative Q-learning, which enables joint decision-making for jamming channels in multi-agent adversarial environments.
[0022] Figure 1 This is a model diagram of an interference system. In this model, a pair of transmitters and receivers constitutes a user pair. N user pairs can communicate simultaneously. User pairs cooperate to determine the communication channel to avoid internal interference between user pairs. The system contains M jammers that interfere with user communications. The jammers have sensing and learning capabilities, able to sense the user's current communication frequency, learn the frequency usage patterns of communication users through intelligent learning algorithms, and generate efficient intelligent jamming strategies to effectively interfere with the communication.
[0023] Figure 2This is a framework diagram of a cooperative Q-learning intelligent jamming method. Each jammer updates its independent Q-value table based on the perceived state and the decisions it makes. Based on the independent Q-value tables maintained by all jammers, the central server of the intelligent jamming system updates the joint Q-value table, which is jointly maintained by all jammers. The central server then makes a joint action a = {a1, ..., a...} based on the joint Q-value table to determine the current state. M This enables distributed computing and collaborative decision-making.
[0024] This invention aims to select the optimal joint jamming channel and utilizes a reinforcement learning algorithm to enable the jammer to interact with the environment to find the best joint jamming strategy. The intelligent jamming method based on cooperative Q-learning proposed in this invention includes the following steps:
[0025] Step 1: Consider an interference scenario consisting of M intelligent jammers and N pairs of communication users (transmit / receive pairs). In this scenario, the communication users cooperate to determine the communication channel to avoid internal interference between them. The users communicate using either a fixed-sequence frequency hopping method or a random-sequence frequency hopping method. The intelligent jammers have sensing and learning capabilities and can perceive the real-time environmental spectrum status s. t ;
[0026] Step 2: Each intelligent jammer establishes and maintains two Q-value tables: an independent Q-value table and a joint Q-value table. The jammer takes the sensed user spectrum state as the state input and selects the joint action a = {a1, ..., a2} according to the ε-greedy strategy. M};
[0027] Step 3: Execute the joint action, evaluate the effect of releasing the interference based on the joint action, and obtain the reward value r for each jammer under the current joint action. m (s t ,a m ), and the total gain R of the entire set of disturbances. t (s,a), and sense the current spectrum state s t+1 ;
[0028] Step 4: Update the independent Q-value table and the combined Q-value table based on the reward values received;
[0029] Step 5: Repeat steps 1 through 4 until the specified number of iterations.
[0030] The specific implementation of the present invention is as follows:
[0031] The communication users of this invention communicate using either a fixed sequence frequency hopping method or a random frequency hopping method, specifically:
[0032] Fixed sequence frequency hopping refers to a user's frequency changes being based on a fixed sequence list. Each time slot sequentially selects a frequency for communication;
[0033] Random sequence frequency hopping refers to a method where the user updates the communication frequency based on a fixed sequence list according to the following strategy:
[0034] The nth pair of users chooses to reside on the current communication frequency with probability ε, i.e.: channel n (t+1) = channel n (t), with probability 1-ε, chooses to jump to the next frequency point, i.e.: channel n (t+1)=[channel n [(t)+1]modK, and the m-th pair of users and the n-th pair of users are at the same time, satisfying channel m (t)≠channel n (t), where t is time.
[0035] The intelligent jammer of this invention can sense the environmental spectrum status in real time. t Specifically:
[0036] The environmental state in which the jammer operates is closely related to the user's current communication channel; therefore, the environmental state space is defined as follows:
[0037] S={s t :s t =(u1(t),…,u n (t))} (1)
[0038] Among them, u n (t)∈[f1,f2,…,f K ], n=1,...,N represents the channel through which the nth pair of communication users communicate at the current time t.
[0039] In this invention, each intelligent jammer establishes and maintains two Q-value tables: an independent Q-value table and a joint Q-value table. The jammer takes the sensed user spectrum state as the state input and selects the joint action a = {a1,...,a2} according to an ε-greedy strategy. M Specifically:
[0040] Q m (s t a) represents the interfering machine j in the independent Q-value table. m In state s t The state-action value, Q(s), for executing the combined action a is: t a) indicates that the set of disturbances in the joint Q-value table is in state s. t The state-action value of the combined action 'a' is related as follows:
[0041]
[0042] Where s t This indicates the current state sensed by the jammer, and 'a' represents the joint action.
[0043] Based on the current perceived state s t jammer j m According to the formula with probability 1-ε Choose a combined action, where a * Indicates state action value The maximum combined interference action is selected; otherwise, a random action is selected. Indicates jammer j m The action space; where the value of ε is continuously updated according to the number of iterations, and the update formula is as follows:
[0044] ε=ε0e -λt (ε0>0,λ>0) (3)
[0045] Where ε0 is the initial value and λ represents the fading coefficient.
[0046] This invention evaluates the effect of releasing interference based on the joint action, and obtains the reward value r of each jammer under the current joint action. m (s t ,a m ), and the total gain R of the entire set of disturbances. t (s,a), specifically:
[0047] Consider quantifying the jamming suppression effect as a benefit value. When the intelligent jammer j m The interference action made a m Capable of successfully suppressing any user channel, the jammer j m The independent payoff is 1, otherwise it is 0; considering cooperation between intelligent jammers, when different intelligent jammers perform the same action, the payoff is... The jammer at time t m The joint benefit is defined as:
[0048]
[0049] Where a m and a n They represent jammer j respectively m and j n The interference decision is the interference channel, u i (t) represents the communication channel of the i-th user at time slot t. δ(·) is the indicator function, and its specific definition is as follows:
[0050]
[0051] For any two values p and q, the value of δ(p,q) is 1 when p and q are equal, and the value of δ(p,q) is 0 when p and q are not equal.
[0052] Different jammers take joint action a={a1,...,a M At that time, the instantaneous reward value and reward sum of each jammer can be obtained. State s t The following joint action a = {a1,...,a2} is executed. M The total gain of the interference set is represented as follows:
[0053]
[0054] This invention updates the independent Q-value table and the joint Q-value table based on the received reward value, specifically as follows:
[0055] jammer j m Update your Q-value table using the following formula:
[0056] Q m (s t ,a t )=(1-α)Q m (s t ,a t )+α[r m (s t ,a m )+γQ m (s t+1 ,a * (7)
[0057] Where α represents the learning rate of the jammer, γ represents the discount factor corresponding to the Q-value update, and s t+1 Represents state s t Execute the joint action a t The next state after that, r m (s t ,a m ) represents the disturbance ensemble in state s t Joint action under the conditions a t Give jammer J m Instant rewards, a * Represents state s t+1 The following joint action maximizes the gain of all intelligent jammers, and this joint action is given by the following formula:
[0058]
[0059] The joint Q-value table is updated according to the following formula:
[0060]
[0061] Example 1
[0062] The first embodiment of the present invention is described in detail below. The system simulation is performed using MATLAB language, and the parameter settings do not affect the generality. This embodiment verifies the effectiveness of the proposed method. Figure 3 , Figure 4 The effectiveness of the fixed-sequence frequency hopping method against users was verified. The parameters were set to consider a system with two intelligent jammers and two user pairs (M=N=2), and both jammers and users had the same number of available channels, 10 channels each (K=10). The user pairs communicated using a fixed-sequence frequency hopping mode, with a hopping period of 0.95ms. The jamming release time slot was set to 0.9ms, the jamming detection time slot to 0.03ms, and the jamming learning time slot to 0.02ms.
[0063] Figure 3 This is a schematic diagram comparing the interference success probability performance of the method proposed in Embodiment 1 of this invention and the comparison algorithm. Figure 4 This is a schematic diagram comparing the user normalized throughput performance of the method proposed in Embodiment 1 of this invention with the comparison algorithm. The comparison algorithm is independent Q-learning, performing calculations every 20 communication time slots. After 50 independent runs, the result is obtained by averaging the results. Figure 3 As can be seen from the interference success probability curve, over time, the interference success rate of the jammer using the cooperative Q-learning jamming method can reach 100%, while the interference success rate of the independent Q-learning algorithm only reaches 50%. Figure 4 The normalized user throughput curve shows that the throughput of the independent Q-learning jamming algorithm ultimately remains around 30%. This is because there is no cooperation between the jammers; each jammer chooses its channel independently. Different jammers can make the same decision at the same time, resulting in wasted jamming resources. The jamming method based on cooperative Q-learning considers the coordination between users and makes the optimal decision to successfully jam two user channels simultaneously. The normalized user throughput gradually decreases and eventually converges, with fluctuations around 5%.
[0064] Example 2
[0065] The second embodiment of the present invention is described in detail below. The system simulation is performed using MATLAB language, and the parameter settings do not affect the generality. This embodiment verifies the effectiveness of the proposed method. Figure 5 , Figure 6The effectiveness of countering user random sequence frequency hopping was verified. The parameters were set as follows: a system with two intelligent jammers and two user pairs (M=N=2), and both jammers and users had the same number of available channels, 10 channels each (K=10). The user pairs communicated using random sequence frequency hopping, with the following rules: users had a 30% probability of choosing to remain on the current communication channel and a 70% probability of hopping to the next channel. The user hopping period was set to 0.95ms, the jamming release time slot to 0.9ms, the jamming sensing time slot to 0.03ms, and the jamming learning time slot to 0.02ms.
[0066] Figure 5 This is a schematic diagram comparing the interference success probability performance of the method proposed in Embodiment 2 of this invention and the comparison algorithm. Figure 6 This is a schematic diagram comparing the user normalized throughput performance of the method proposed in Embodiment 2 of the present invention with that of the comparison algorithm. The comparison algorithm is independent Q-learning, performing calculations every 20 communication time slots. The result is obtained by averaging the results after 50 independent runs. Figure 5 The interference success probability curves show that when the jammer uses the cooperative Q-learning algorithm, it can interfere with the communication channel with a certain probability. However, when the jammer uses the independent Q-learning algorithm, the success rate is lower due to the uncertainty of user channel switching and the independence between jammers. With a user switching probability of 70%, the cooperative Q-learning algorithm can successfully interfere with the channel with a 70% probability. Figure 6 The normalized user throughput curve shows that when the jammer uses the independent Q-learning algorithm, approximately 60% of the data can be transmitted normally, while 40% of the user data is successfully blocked. When the jammer uses a jamming method based on cooperative Q-learning, approximately 35% of the data can be transmitted normally, while 65% of the user data is successfully blocked. Figure 4 The large fluctuations in the curve are due to the uncertainty of user channel switching. When counting every 20 time slots, the number of times a channel is selected for camping is uncertain. When a user chooses to camp, the jammer tends to select the next channel with the larger Q value, which can lead to decision errors at this point, hence the curve fluctuations.
[0067] Comparative analysis revealed that the interference method based on cooperative Q-learning proposed in this invention can effectively interfere with user communications, greatly improving the interference effect.
[0068] In summary, the interference method based on cooperative Q-learning proposed in this invention can coordinate the internal decision-making of a multi-agent jammer, effectively improving the interference effect in multi-agent adversarial scenarios. The jammer does not require prior user or channel information during the decision-making process; it can find the optimal channel decision simply by interacting with the spectrum environment.
Claims
1. A smart interference method based on cooperative Q-learning, characterized in that, Includes the following steps: Step 1, consider from A smart jammer and For interference scenarios composed of communication users, and Let represent the number of intelligent jammers and the number of communication users, respectively, satisfying . In jamming scenarios, communication users determine communication channels through cooperation. They communicate using either fixed-sequence frequency hopping or random-sequence frequency hopping. The intelligent jammer possesses sensing and learning capabilities, enabling it to perceive the real-time environmental spectrum status. ,in Indicates the current moment. A time-frequency diagram of the environment can characterize the distribution of signals in time and frequency. Step 2: Each intelligent jammer establishes and maintains an independent Q-value table and a joint Q-value table. Q-learning is a reinforcement learning algorithm based on a value function. The Q-value table records the future state-action values of state-action pairs in Q-learning. The independent Q-value table records the future state-action values of environmental spectrum state-action pairs for each jammer. The joint Q-value table contains the state-action values from all the independent Q-value tables of the jammers. The jammer takes the perceived environmental spectrum state as its state input and, according to... - Greedy strategy selects combined interference actions That is, the set of all jammer jamming channels, where Indicates jammer Interference channels, Indicates the first One jammer, satisfying ; Step 3: Execute the joint jamming action, evaluate the effect of releasing the jamming based on the joint jamming action, and obtain the expected value of the revenue of each jammer under the current joint jamming action. And the overall profit value of the entire jamming machine set. And sense and obtain the current spectrum state. Among them, the profit value Used to evaluate actions In the environmental spectrum state Whether it is good or bad, Used to evaluate combined interference actions In the environmental spectrum state The quality of the rain; Step 4: Update the independent Q-value table and the joint Q-value table based on the earned revenue. Step 5: Repeat steps 1 through 4 until the specified number of iterations is reached.
2. The intelligent interference method based on cooperative Q-learning according to claim 1, characterized in that, In step 1, the communication users use either a fixed sequence frequency hopping method or a random frequency hopping method for communication, specifically: Fixed sequence frequency hopping refers to a user's communication channel being based on a fixed sequence list. The change occurs when each time slot user sequentially selects a channel for communication. This represents the total number of available channels in the sequence list. Indicates the first in the sequence list One channel, satisfying ; Random sequence frequency hopping refers to a method where users update the communication channel based on a fixed sequence list, following this strategy: No. For users with probability Choose to remain on the current communication channel with probability. Choose to switch to the next channel, and the first For users and For users at the same time, different communication channels are required, among which... Let be a real number that satisfies .
3. The intelligent interference method based on cooperative Q-learning according to claim 2, characterized in that, In step 1, the intelligent jammer can sense the environmental spectrum status in real time. Specifically: The environmental conditions in which the jammer operates are closely related to the user's current communication channel; therefore, the environmental spectrum conditions... The definition is as follows: (1) in, Indicates the first For communication users at any time The communication channel satisfies .
4. The intelligent interference method based on cooperative Q-learning according to claim 3, characterized in that, In step 2, each intelligent jammer establishes and maintains an independent Q-value table and a joint Q-value table. The jammer will then process the environmental spectrum state it senses. As a state input, according to - Greedy strategy selects combined interference actions Specifically: In the table of independent Q values, the jammer is represented. In the environmental spectrum state Execute joint interference actions State action value, This indicates that the set of jammers in the joint Q-value table is in state. Execute joint interference actions The relationship between the state and action values is as follows: (2) in express The environmental spectrum state at any given time. Indicates a combined interference action; Based on the current environmental spectrum status jammer With probability According to the formula Select a combined interference action, where Indicates state action value The maximum combined interference action is selected; otherwise, a random action is selected. , Indicates jammer The action space, i.e., the set of currently available interference channels; where The value is updated continuously based on the number of iterations, and the update formula is as follows: (3) in As the initial value, Let represent the fading coefficient, and satisfy . .
5. The intelligent interference method based on cooperative Q-learning according to claim 4, characterized in that, In step 3, the effectiveness of releasing the jamming is evaluated based on the joint jamming action, and the benefit value of each jammer under the current joint jamming action is obtained. And the overall profit value of the entire jamming machine set. Specifically: Consider quantifying the jamming suppression effect as a benefit value, when the intelligent jammer Interference actions performed Capable of successfully suppressing any user channel, the jammer The independent payoff is 1, otherwise it is 0; considering cooperation between intelligent jammers, when different intelligent jammers perform the same action, the payoff is... ;Will Time jammer The revenue is defined as: (4) in and They represent jammers and Interference decision channel, This indicates a time slot. Time Communication channels for individual users; This is an indicator function, and its specific definition is as follows: (5) For any two values and ,when and When they are equal, The value is 1, when and When they are not equal The value is 0; Different jammers take joint jamming actions At that time, the instantaneous benefit value of each jammer can be obtained; State Execute joint interference actions The total profit of the jammer set is expressed as follows: (6)。 6. The intelligent interference method based on cooperative Q-learning according to claim 5, characterized in that, In step 4, the independent Q-value table and the joint Q-value table are updated based on the received revenue, specifically as follows: jammer Update your Q-value table using the following formula: (7) in, This indicates the learning rate of the jammer. This represents the discount factor corresponding to the Q-value update; this parameter determines the importance of future returns in the current decision. Representing state Execute joint interference actions The next state after that, Indicates the jammer set in state Combined interference actions under the conditions give jammer The profit value, Representing state The following joint jamming action maximizes the benefit of all intelligent jammers, and is given by the following formula: (8) The joint Q-value table is updated according to the following formula: (9)。
Citation Information
Patent Citations
Deep Q neural network anti-interference model and intelligent anti-interference algorithm
CN108777872A
Multi-agent intelligent electronic interference method based on information sharing
CN113049885A