Two-Dimensional Intelligent Anti-Jamming Decision-Making Method and System Based on Q-Learning Algorithm

Through the two-dimensional intelligent anti-interference decision-making method based on Q learning algorithm, the problem of inability to adapt to complex environment changes in the existing technology is solved, efficient anti-interference decision-making under unknown jammer information is achieved, and communication reliability and real-time performance are improved.

CN116390259BActive Publication Date: 2025-07-25XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310448479.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-24
Publication Date
2025-07-25
Estimated Expiration
2043-04-24

AI Technical Summary

Technical Problem

The existing wireless communication anti-interference technology is mostly blind anti-interference method, which cannot adapt to complex environment changes in time, and relies on manual experience, resulting in communication interruptions and long reaction time.

Method used

A two-dimensional intelligent anti-interference decision-making method based on Q learning algorithm is adopted. By dividing time intervals into discrete time slots, the transmitter and the jammer play in each time slot, and the ε-greedy algorithm is used to update the value function of the state-action pair to obtain the optimal anti-interference strategy.

Benefits of technology

In the case of unknown jammer information, anti-interference performance is improved, which is better than traditional random selection strategies, and more intelligent anti-interference decisions are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116390259B_ABST
    Figure CN116390259B_ABST
Patent Text Reader

Abstract

A two-dimensional intelligent anti-jamming decision-making method and system based on the Q-learning algorithm. The method includes equally dividing continuous time into a number of discrete time slots, conducting a round of game in each time slot, and dividing each time slot into two parts, Δt1 and Δt2. In the Δt1 part, the transmitter and the receiver communicate according to the selected communication scheme, and the jammer jams according to the selected jamming scheme. In the Δt2 part, the signal-to-interference-plus-noise ratio at the receiving end is fed back to the transmitter, and the transmitter calculates the utility of the current time slot and updates the transmitter time slot state and the state-action pair value function. The transmitter selects the communication scheme for the next time slot according to the value function of the state-action pair and the ε-greedy algorithm, and the jammer selects the jamming scheme for the next time slot through the same method. By repeating the above process, the optimal value function of the transmitter state-action pair is obtained. The present invention can make decisions without knowing the relevant information of the jammer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of anti - interference in wireless communication, and particularly relates to a two - dimensional intelligent anti - interference decision - making method and system based on the Q - learning algorithm. Background Art

[0002] Due to reasons such as the rapid growth of wireless communication services, the general increase in terminal devices, and the increasingly complex electromagnetic environment, not only is the wireless spectrum resource becoming increasingly tense, but also the interference suffered by wireless communication is becoming increasingly serious. How to conduct efficient and reliable communication in such an environment has been widely concerned and studied. Especially in the military information field, the level of informatization technology is an important aspect of the military strength. In a battlefield environment characterized by information confrontation and network warfare, military anti - interference has become a difficult problem. In order to ensure the real - time and reliability of communication on future battlefields, it is essential to research more intelligent anti - interference communication technologies for complex interference environments. Now, this direction has become an important research topic.

[0003] Generally, communication anti - interference technologies mainly include three categories: one is frequency - domain processing, such as direct - sequence spread - spectrum, frequency - hopping, etc.; the second is spatial - domain processing, such as adaptive antennas, etc.; the third is time - domain processing, such as burst communication. These anti - interference technologies each have their own advantages, but they all belong to the "blind anti - interference method", that is, the anti - interference ability is determined at the beginning of system design. Once the interference from the attacker exceeds its anti - interference tolerance, communication interruption will occur. From the perspective of anti - interference strategies, the currently common decision - making technology is based on manual intervention, passive, and mechanical manual operation. This way of passively adjusting the communication waveform by feedback of communication quality has many defects, such as complex operation, dependence on manual experience, long response time, easy to cause long - term communication interruption, and inability to adapt to environmental changes in a timely manner. In contrast, intelligent anti - interference technology, in essence, introduces the idea of cognitive radio into communication anti - interference technology to achieve a more intelligent anti - interference technology with certain cognitive abilities. And intelligent anti - interference decision - making technology, as one of the cores of intelligent anti - interference technology, is an important link in anti - interference decision - making technology. Summary of the Invention

[0004] The purpose of the present invention is to provide a two - dimensional intelligent anti - interference decision - making method and system based on the Q - learning algorithm for the above - mentioned problems in the prior art, which can make anti - interference decisions when the relevant information of the jammer is unknown and has excellent anti - interference performance.

[0005] To achieve the above purpose, the present invention has the following technical solutions:

[0006] In the first aspect, a two - dimensional intelligent anti - interference decision - making method based on the Q - learning algorithm is provided, including:

[0007] The continuous time is equally divided into a number of discrete time slots, and one round of game is carried out in each time slot. Each time slot is divided into two parts, Δt1 and Δt2;

[0008] In the Δt1 part of each time slot, the transmitter and the receiver communicate according to the selected communication scheme, and the jammer jams the receiver according to the selected jamming scheme;

[0009] In the Δt2 part of each time slot, the transmitter obtains the utility within the current time slot and the communication scheme and updates the state of the transmitter in the current time slot as and updates the value function of the state-action pair of the transmitter The jammer obtains the utility within the current time slot and the jamming scheme and updates the state of the jammer in the current time slot as and updates the value function of the state-action pair of the jammer

[0010] The transmitter selects the communication scheme for the next time slot by using the ε-greedy algorithm according to the value function of the state-action pair of the transmitter The jammer selects the jamming scheme for the next time slot by using the ε-greedy algorithm according to the value function of the state-action pair of the jammer

[0011] Repeat the above process until the value function of the state-action pair of the transmitter converges or reaches the maximum number of iterations, and obtain the optimal value function of the state-action pair of the transmitter

[0012] The transmitter makes a decision according to the state of the current time slot by using the optimal value function of the state-action pair of the transmitter

[0013] As a preferred solution, the communication scheme selected by the transmitter is:

[0014] a s =(P s , C s )

[0015] where P s represents the transmission power selected by the transmitter, C s represents the communication channel selected by the transmitter, and the selected communication channel comes from n c ​Available channels;

[0016] The interference scheme selected by the jammer is:

[0017] a j =(P j , C j )

[0018] Where P j represents the transmission power selected by the jammer, C j represents the communication channel selected by the jammer, and the selected communication channel comes from n c available channels;

[0019] The participants in the game are the transmitter and the jammer;

[0020] The action space of the transmitter is all optional communication schemes a s , and the game payoff is the utility u s between the transmitter and the receiver; the action space of the jammer is all optional interference schemes a j , and the game payoff is the utility u j of the jammer;

[0021] The utility between the transmitter and the receiver is calculated as follows:

[0022]

[0023] Where P s represents the transmission power of the transmitter; h s represents the channel gain between the transmitter and the receiver; P j represents the transmission power of the jammer; h j represents the channel gain between the jammer and the receiver; l s represents the utility loss generated by the transmission power of the transmitter; σ 2 represents the noise power of the wireless channel; f(ζ) is an indicator function,

[0024] The jammer reduces the communication quality between the transmitter and the receiver by interfering with the receiver, and the utility of the jammer is:

[0025]

[0026] Where l j represents the utility loss generated by the transmission power of the jammer.

[0027] As a preferred solution, the receiver feeds back the signal-to-interference-plus-noise ratio (SINR) information at the receiving end to the transmitter through a feedback channel, and the transmitter makes a decision based on the SINR information at the receiving end; taking the maximization of the SINR at the receiving end as the optimization objective, the mixed-strategy Nash equilibrium in the game process between the transmitter and the jammer is solved to obtain the optimal anti-jamming strategy of the transmitter and the optimal jamming strategy of the jammer.

[0028] As a preferred solution, the anti-jamming confrontation information expression of the transmitter in the k-th time slot is:

[0029]

[0030] In the formula, is the utility of the transmitter in the k-th time slot; is the communication scheme selected by the transmitter in the k-th time slot;

[0031] The state of the transmitter in the k-th time slot is:

[0032]

[0033] In the formula, w represents the number of confrontation information included in each time slot state, that is, the time backtracking length of the confrontation information.

[0034] As a preferred solution, the jamming confrontation information expression of the jammer in the k-th time slot is:

[0035]

[0036] In the formula, is the utility of the jammer in the k-th time slot; is the jamming scheme selected by the jammer in the k-th time slot;

[0037] The state of the jammer in the k-th time slot is composed of the jamming confrontation information of the previous w time slots, and the time slot state is as follows:

[0038]

[0039] where w represents the number of confrontation information included in each time slot state, that is, the time backtracking length of the confrontation information.

[0040] As a preferred solution, update the value function of the transmitter state-action pair The expression is as follows:

[0041]

[0042] Among them, γ represents the discount factor of the long-term return;

[0043] The transmitter is based on the value function of the transmitter state-action pair Use the ε-greedy algorithm to select the communication scheme of the transmitter for the next time slot When doing so, the transmitter selects the communication scheme with a probability of 1 - ε where satisfy with probability to select the remaining communication schemes; where n represents the number of the remaining communication schemes

[0044] In a second aspect, a two-dimensional intelligent anti-jamming decision-making system based on the Q-learning algorithm is provided, including:

[0045] A time slot division module, configured to equally divide continuous time into a plurality of discrete time slots, conduct a round of game in each time slot, and divide each time slot into two parts, Δt1 and Δt2

[0046] An execution module, configured to, in the Δt1 part of each time slot, the transmitter and the receiver communicate according to the selected communication scheme, and the jammer jams the receiver according to the selected jamming scheme

[0047] An update module, configured to, in the Δt2 part of each time slot, the transmitter obtains the utility within the current time slot and the communication scheme update the state of the transmitter for the current time slot as and update the value function of the state-action pair of the transmitter The jammer obtains the utility within the current time slot and the jamming scheme update the state of the jammer for the current time slot as and update the value function of the state-action pair of the jammer

[0048] A next time slot scheme selection module, configured to, for the transmitter, according to the value function of the state-action pair of the transmitter use the ε-greedy algorithm to select the communication scheme of the transmitter for the next time slot For the jammer, according to the value function of the state-action pair of the jammer use the ε-greedy algorithm to select the jamming scheme of the jammer for the next time slot

[0049] An optimal value function acquisition module, configured to repeat the above process until the value function of the state-action pair of the transmitter converges or reaches the maximum number of iterations, and obtain the optimal value function of the state-action pair of the transmitter

[0050]

[0051] The optimal value function decision module is used for the transmitter to make decisions according to the current time slot state using the optimal value function of the transmitter state-action pair to make decisions.

[0052] In a third aspect, an electronic device is provided, including:

[0053] a memory storing at least one instruction; and

[0054] a processor that executes the instructions stored in the memory to implement the two-dimensional intelligent anti-jamming decision method based on the Q-learning algorithm.

[0055] In a fourth aspect, a computer-readable storage medium is provided, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the two-dimensional intelligent anti-jamming decision method based on the Q-learning algorithm.

[0056] Compared with the prior art, the present invention has at least the following beneficial effects:

[0057] By using game theory to establish the confrontation relationship between the jammer and the transmitter, a Q-learning network is established at the transmitter. The network input is the transmitter utility and the selected communication scheme in the previous few time slots, and the network output is the communication scheme selected by the transmitter in the next time slot. At the end of each time slot, the transmitter will obtain the utility and communication scheme in this time slot, update the time slot state and the state-action pair value function, and then select the communication scheme of the transmitter in the next time slot according to the state-action pair value function. Repeat the above process, update the state-action pair value function to obtain the optimal value function, and then obtain the optimal anti-jamming strategy of the transmitter through the optimal value function. When the jammer does not know the transmitter information, the jammer's interference decision process based on the Q-learning network is similar to the transmitter's anti-jamming decision process. The two-dimensional intelligent anti-jamming decision method based on the Q-learning algorithm of the present invention can make decisions without knowing the relevant information of the jammer, and its anti-jamming performance is better than the traditional randomly selected anti-jamming strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and those of ordinary skill in the art can obtain other related drawings without creative efforts based on these drawings.

[0059] Figure 1 Schematic diagram of the working scenario of the transmitter in the embodiment of the present invention;

[0060] Figure 2Schematic diagram of time slot division in an embodiment of the present invention;

[0061] Figure 3 Graph showing the comparison results of the anti - interference performance in an embodiment of the present invention and that of traditional anti - interference strategies. Detailed implementation manners

[0062] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, those of ordinary skill in the art can also obtain other embodiments without making creative efforts.

[0063] As Figure 1 shown, in the scenario of an embodiment of the present invention, there is a pair of a transmitter and a receiver communicating. There is a jammer in the environment that emits interference signals to interfere with the normal communication between the transmitter and the receiver by interfering with the receiver's reception of the signals. The distance between the transmitter and the receiver is d s , and the distance between the jammer and the receiver is d j . The optional transmission power of the transmitter is n sp represents the total number of optional discrete powers of the transmitter. The optional transmission power of the jammer is n jp represents the total number of optional discrete powers of the jammer. There are n c available channels between the transmitter and the receiver.

[0064] The communication scheme selected by the transmitter is:

[0065] a s =(P s , C s ) (1)

[0066] where P s represents the transmission power selected by the transmitter, C s represents the communication channel selected by the transmitter, and the selected channel comes from n c available channels.

[0067] The interference scheme selected by the jammer is:

[0068] a j =(P j , C j ) (2)

[0069] where P j represents the transmission power selected by the jammer, C jDenote the communication channel selected by the jammer, and the selected channel is from n c available channels.

[0070] Each time slot is a game round. As Figure 2 shown, each time slot is divided into two parts, Δt1 and Δt2. At the beginning of each time slot, i.e., within Δt1, the transmitter and the jammer respectively execute their communication and jamming schemes; at the end of each time slot, i.e., within Δt2, the transmitter and the jammer select their communication and jamming schemes for the next time slot according to their respective strategies. Within the same time slot, the schemes of the transmitter and the jammer remain unchanged.

[0071] To maximize the signal-to-interference-plus-noise ratio (SINR) at the receiver and at the same time consider the power cost generated by the transmit power selected by the transmitter, the utilities between the transmitter and the receiver are as follows:

[0072]

[0073] where P s denotes the transmit power of the transmitter; h s denotes the channel gain between the transmitter and the receiver; P j denotes the transmit power of the jammer; h j denotes the channel gain between the jammer and the receiver; l s denotes the utility loss generated by the transmit power of the transmitter; σ 2 denotes the noise power of the wireless channel; f(ζ) is an indicator function, defined as follows:

[0074]

[0075] The jammer reduces the communication quality between the transmitter and the receiver by jamming the receiver.

[0076] The utility of the jammer is designed as follows:

[0077]

[0078] where l j denotes the utility loss generated by the transmit power of the jammer.

[0079] The receiver feeds back the SINR information at the receiving end to the transmitter through a feedback channel. The transmitter makes decisions based on the feedback information.

[0080] Each time slot is regarded as a game round. The transmitter and the jammer simultaneously select their communication and jamming schemes for the next time slot at the end of the time slot. This problem is modeled as a strategic-form game problem, and the players of the game are the transmitter and the jammer. The action space of the transmitter is all optional communication schemes a s, the game payoff is the transmitter utility \(u\) s ; the action space of the jammer is all optional jamming schemes \(a\) j , the game payoff is the jammer utility \(u\) j .

[0081] According to the relevant theories of game theory, there exists a mixed-strategy Nash equilibrium in this game. When the transmitter and the jammer can obtain perfect information about each other, the Lemke-Howson algorithm is used to solve the Nash equilibrium, and the optimal anti-jamming strategy of the transmitter and the optimal jamming strategy of the jammer can be obtained.

[0082] However, in the actual process, it is difficult for the transmitter to obtain perfect information about the jammer. In the case where the transmitter does not know the information of the jammer, the embodiments of the present invention use the Q-learning algorithm to solve the optimal anti-jamming strategy of the transmitter.

[0083] First, assume that the jammer can obtain all the information of the transmitter and the receiver, take the utility as the game payoff, and obtain the optimal jamming strategy of the jammer from the mixed-strategy Nash equilibrium.

[0084] The anti-jamming confrontation information of the transmitter at the \(k\)th time slot is:

[0085]

[0086] where is the utility of the transmitter at the \(k\)th time slot; is the communication scheme selected by the transmitter within the \(k\)th time slot.

[0087] The state of the transmitter at the \(k\)th time slot is:

[0088]

[0089] where \(w\) represents the number of confrontation information included in the state of each time slot, that is, the time backtracking length of the confrontation information.

[0090] Within \(\Delta t_1\) of each time slot, the transmitter and the receiver communicate according to the selected communication scheme, and the jammer jams the receiver according to the selected jamming scheme.

[0091] Within \(\Delta t_2\) of each time slot, the receiver feeds back the signal-to-interference-plus-noise ratio at the receiving end to the transmitter through the feedback channel, and the transmitter obtains the utility and the communication scheme The time slot state is updated from to Update the state-action pair value function of the transmitter using the following formula

[0092]

[0093] Among them, γ represents the discount factor of the long-term return.

[0094] The transmitter selects the communication scheme between the transmitter and the receiver in the next time slot according to the value function and the ε-greedy algorithm The transmitter selects the communication scheme with a probability of 1 - ε Among them, satisfies With probability, the remaining communication schemes are selected. Among them, n represents the number of the remaining communication schemes.

[0095] Select the communication scheme in the next time slot according to the above principle Use this communication scheme within Δt1 of the next time slot.

[0096] Repeat the above process until the value function converges or reaches the maximum number of iterations to obtain the optimal value function

[0097]

[0098] Within Δt2 of each time slot, obtain the utility within the current time slot and the communication scheme Update the time slot state s k-1 to s k . According to the optimal value function select the communication scheme that satisfies as the communication scheme in the next time slot. as the communication scheme in the next time slot.

[0099] However, for the jammer, it is also difficult to obtain all the information of the transmitter and the receiver. Therefore, a Q-learning network is also established at the jammer, and the optimal strategy of the jammer is obtained by learning the historical experience of the jammer. It is assumed that the jammer can steal the signal-to-interference-plus-noise ratio information fed back by the receiver to the transmitter to enhance the performance of the jammer and test the anti-jamming ability of the scheme proposed in the embodiment of the present invention.

[0100] The interference countermeasure information of the jammer in the kth time slot is:

[0101]

[0102] Among them, is the utility of the jammer in the kth time slot; is the interference scheme selected by the jammer in the kth time slot.

[0103] The state of the jammer in the kth time slot is composed of the interference countermeasure information of the previous w time slots, and the time slot state is as follows:

[0104]

[0105] Among them, w represents the number of adversarial information included in each time slot state, that is, the time backtracking length of the adversarial information.

[0106] Within Δt1 of each time slot, the transmitter and the receiver communicate according to the selected communication scheme, and the jammer jams the receiver according to the selected jamming scheme.

[0107] Within Δt2 of each time slot, the transmitter obtains the utility within the current time slot and the communication scheme updates the time slot state for updating the value function of the state-action pair Meanwhile, the jammer obtains the utility within the current time slot and the jamming scheme updates the time slot state for updating the value function of the state-action pair

[0108] Then, the transmitter and the jammer respectively use their respective value functions and to select the communication scheme of the transmitter in the next time slot using the ε-greedy algorithm and the jamming scheme of the jammer

[0109] Repeat the above process until the value function converges or reaches the maximum number of iterations, and obtain the optimal value function

[0110]

[0111] After obtaining the optimal value function, the transmitter makes a decision according to the current time slot state using the optimal value function to make a decision and select an action that satisfies as the communication scheme for the next game round.

[0112] To verify the performance of the anti-jamming decision-making method of the embodiment of the present invention, the following simulations are carried out:

[0113] The number of available communication channels for the transmitter and the receiver is 3. The optional transmission powers of the transmitter are {1W, 2W, 3W}, and the optional transmission powers of the jammer are {3W, 5W}. The distance between the transmitter and the receiver is 1000m, and the distance between the jammer and the receiver is 600m. The noise power of the wireless channel is -114dBm.

[0114] In the embodiments of the present invention, there are two cases for the jammer, namely, the jammer can obtain all the information of the transmitter and the jammer cannot obtain all the information of the transmitter. The above two cases are distinguished by using "perfect information" and "imperfect information" as identifiers.

[0115] There are also two comparison schemes as follows:

[0116] 1) Scheme 1: The transmitter can obtain all the information of the jammer, and the utility of the transmitter is used as the game payoff; the jammer can obtain all the information of the transmitter, and the utility of the jammer is used as the game payoff. The mixed strategy Nash equilibrium is used as the optimal anti-jamming strategy of the transmitter and the optimal jamming strategy of the jammer;

[0117] 2) Scheme 2: The transmitter does not have the relevant information of the jammer, and the transmitter adopts a traditional random selection strategy to make anti-jamming decisions. The jammer cannot obtain all the information of the transmitter, and a Q-learning network is set at the jammer to obtain the optimal jamming strategy of the jammer.

[0118] Figure 3 The comparison results of the four schemes are shown. In comparison scheme 1, the transmitter has all the information of the jammer, and the utility of the transmitter is 4.10. The scheme proposed in the embodiments of the present invention is carried out under the condition of unknown jammer information. Due to the loss of relevant information of the jammer, the utility has decreased. In the case of perfect information, the utility of the embodiments of the present invention reaches 3.74, which is 91% of the utility in comparison scheme 1. In the case of imperfect information, since the strategy of the jammer is time-varying, which increases the uncertainty of the environment, the utility of the transmitter has decreased, but it also reaches 3.36. The utility of comparison scheme 2 is 2.74, and the utility of the scheme proposed in the embodiments of the present invention is 22% higher than that of scheme 2 in the case of imperfect information.

[0119] Therefore, in summary, it can be seen that the two-dimensional anti-jamming decision-making method based on the Q-learning algorithm proposed in the embodiments of the present invention can make anti-jamming decisions when the relevant information of the jammer is unknown, and the anti-jamming performance is better than that of the traditional random selection strategy.

[0120] Another embodiment of the present invention also proposes a two-dimensional intelligent anti-jamming decision-making system based on the Q-learning algorithm, including:

[0121] A time slot division module, configured to equally divide continuous time into a plurality of discrete time slots, perform a round of game in each time slot, and divide each time slot into two parts, Δt1 and Δt2;

[0122] An execution module, configured to, in the Δt1 part of each time slot, the transmitter and the receiver communicate according to the selected communication scheme, and the jammer jams the receiver according to the selected jamming scheme;

[0123] An update module, which is used for the transmitter to obtain the utility within the current time slot during the Δt2 part of each time slot and the communication scheme Update the current time slot state of the transmitter For And update the value function of the transmitter state-action pair The jammer obtains the utility within the current time slot and the jamming scheme Update the current time slot state of the jammer For And update the value function of the jammer state-action pair

[0124] The next time slot scheme selection module is used for the transmitter to select the communication scheme of the transmitter in the next time slot according to the value function of the transmitter state-action pair Using the ε-greedy algorithm to select the communication scheme of the transmitter in the next time slot The jammer selects the jamming scheme of the jammer in the next time slot according to the value function of the jammer state-action pair Using the ε-greedy algorithm to select the jamming scheme of the jammer in the next time slot

[0125] The optimal value function acquisition module is used to repeat the above process until the value function of the transmitter state-action pair Converges or reaches the maximum number of iterations, and obtains the optimal value function of the transmitter state-action pair

[0126]

[0127] The optimal value function decision module is used for the transmitter to make a decision according to the current time slot state Using the optimal value function of the transmitter state-action pair To make a decision.

[0128] In the embodiment of the present invention, for a wireless environment with an intelligent jammer, the wireless anti-jamming problem is modeled as a strategic game problem. When the two adversarial parties cannot obtain each other's information, the two adversarial parties respectively design an adversarial learning method based on the Q-learning algorithm, so as to obtain the interference countermeasure decision under unknown information.

[0129] Another embodiment of the present invention also proposes an electronic device, including: a memory that stores at least one instruction; and a processor that executes the instruction stored in the memory to implement the two-dimensional intelligent anti-jamming decision method based on the Q-learning algorithm.

[0130] Another embodiment of the present invention further provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the two-dimensional intelligent anti-interference decision-making method based on the Q-learning algorithm.

[0131] Exemplarily, the instructions stored in the memory can be divided into one or more modules / units, which are stored in the computer-readable storage medium and executed by the processor to complete the two-dimensional intelligent anti-interference decision-making method based on the Q-learning algorithm of the present invention. The one or more modules / units can be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the server.

[0132] The electronic device can be a computing device such as a smart phone, a notebook, a palm computer, and a cloud server. The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the electronic device may further include more or fewer components, or combine certain components, or different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.

[0133] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0134] The memory may be an internal storage unit of the server, such as the hard disk or memory of the server. The memory may also be an external storage device of the server, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the server. Further, the memory may also include both the internal storage unit and the external storage device of the server. The memory is used to store the computer-readable instructions and other programs and data required by the server. The memory may also be used to temporarily store the data that has been output or will be output.

[0135] It should be noted that for the content such as information interaction and execution process between the above-mentioned module units, since it is based on the same concept as the method embodiment, for its specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be repeated here.

[0136] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment, and details will not be repeated here.

[0137] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc.

[0138] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0139] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A two-dimensional intelligent anti-interference decision-making method based on the Q-learning algorithm, characterized in that It includes: The continuous time is equally divided into a number of discrete time slots, and one round of game is carried out in each time slot. Each time slot is divided into two parts, Δt1 and Δt2. In the Δt1 part of each time slot, the transmitter and the receiver communicate according to the selected communication scheme, and the jammer jams the receiver according to the selected jamming scheme. During the Δt2 portion of each time slot, the transmitter obtains the utility within the current time slot and the communication scheme Update the current time slot state of the transmitter For and update the value function of the transmitter state-action pair The jammer obtains the utility within the current time slot and the jamming scheme Update the current time slot state of the jammer For and update the value function of the jammer state-action pair The transmitter selects the communication scheme of the transmitter in the next time slot by using the ε-greedy algorithm according to the value function of the transmitter state-action pair The transmitter selects the communication scheme of the transmitter in the next time slot by using the ε-greedy algorithm The jammer selects the jamming scheme of the jammer in the next time slot by using the ε-greedy algorithm according to the value function of the jammer state-action pair The jammer selects the jamming scheme of the jammer in the next time slot by using the ε-greedy algorithm Repeat the above process until the value function of the transmitter state-action pair converges or reaches the maximum number of iterations, and obtain the optimal value function of the transmitter state-action pair The transmitter makes a decision based on the current time slot status using the optimal value function of the transmitter state-action pair to make a decision.

2. The two-dimensional intelligent anti-interference decision-making method based on the Q-learning algorithm according to claim 1, characterized in that The communication scheme selected by the transmitter is: a s = (P s , C s ) Where, P s represents the transmission power selected by the transmitter, C s represents the communication channel selected by the transmitter, and the selected communication channel is from n c available channels; The jamming scheme selected by the jammer is: a j =(P j ,C j ) where P j represents the transmit power selected by the jammer, C j represents the communication channel selected by the jammer, and the selected communication channel is from n c available channels; The participants in the game are the transmitter and the jammer. The action space of the transmitter is all the optional communication schemes a s , and the game payoff is the utility u between the transmitter and the receiver s ; The action space of the jammer is all the optional jamming schemes a j , and the game payoff is the utility u of the jammer j ; The utility between the transmitter and the receiver is calculated by the following formula: Wherein, P s represents the transmission power of the transmitter; h s represents the channel gain between the transmitter and the receiver; P j represents the transmission power of the jammer; h j represents the channel gain between the jammer and the receiver; l s represents the utility loss generated by the transmission power of the transmitter; σ 2 represents the noise power of the wireless channel; f(ζ) is an indicator function, The jammer reduces the communication quality between the transmitter and the receiver by jamming the receiver, and the utility of the jammer is: where l j represents the utility loss caused by the transmitting power of the jammer.

3. The two-dimensional intelligent anti-interference decision-making method based on the Q-learning algorithm according to claim 1, characterized in that The receiver feeds back the signal-to-interference-plus-noise ratio (SINR) information at the receiving end to the transmitter through a feedback channel, and the transmitter makes a decision according to the SINR information at the receiving end. Taking the maximization of the SINR at the receiving end as the optimization objective, the mixed-strategy Nash equilibrium in the game process between the transmitter and the jammer is solved to obtain the optimal anti-jamming strategy of the transmitter and the optimal jamming strategy of the jammer.

4. The two-dimensional intelligent anti-interference decision-making method based on the Q-learning algorithm according to claim 1, characterized in that, The anti-jamming confrontation information expression of the transmitter in the k-th time slot is: wherein, is the utility of the transmitter in the k-th time slot; is the communication scheme selected by the transmitter within the k-th time slot; The state of the transmitter in the k-th time slot is: In the formula, w represents the number of confrontation information included in each time slot state, that is, the time backtracking length of the confrontation information.

5. The two-dimensional intelligent anti-interference decision-making method based on the Q-learning algorithm according to claim 1, characterized in that The jamming confrontation information expression of the jammer in the k-th time slot is: wherein, is the utility of the jammer in the k-th time slot; is the jamming scheme selected by the jammer within the k-th time slot; The state of the jammer in the k-th time slot is composed of the jamming confrontation information of the previous w time slots, and the time slot state is as follows: Among them, w represents the number of confrontation information included in each time slot state, that is, the time backtracking length of the confrontation information.

6. The two-dimensional intelligent anti-interference decision-making method based on the Q-learning algorithm according to claim 1, characterized in that The value function for updating the transmitter state-action pair The expression is as follows: Among them, γ represents the discount factor of the long-term reward; The transmitter selects the communication scheme of the transmitter in the next time slot by using the ε-greedy algorithm according to the value function of the transmitter state-action pair When selecting the communication scheme of the transmitter in the next time slot by using the ε-greedy algorithm the transmitter selects the communication scheme with a probability of 1 - ε where it satisfies and selects the remaining communication schemes with a probability of ; where n represents the number of the remaining communication schemes 7. A two-dimensional intelligent anti-jamming decision-making system based on the Q-learning algorithm, characterized in that, It includes: A time slot division module, which is used to equally divide the continuous time into a number of discrete time slots, carry out one round of game in each time slot, and divide each time slot into two parts, Δt1 and Δt2. An execution module, which is used to, in the Δt1 part of each time slot, enable the transmitter and the receiver to communicate according to the selected communication scheme, and enable the jammer to jam the receiver according to the selected jamming scheme. An update module, which is used for the transmitter to obtain the utility within the current time slot during the Δt2 part of each time slot and the communication scheme Update the current time slot state of the transmitter For And update the value function of the state-action pair of the transmitter The jammer obtains the utility within the current time slot and the jamming scheme Update the current time slot state of the jammer For And update the value function of the state-action pair of the jammer The next time slot scheme selection module is used for the transmitter to select the communication scheme of the next time slot transmitter according to the value function of the transmitter state-action pair using the ε-greedy algorithm The jammer selects the jamming scheme of the next time slot jammer according to the value function of the jammer state-action pair using the ε-greedy algorithm The optimal value function acquisition module is used to repeat the above process until the value function of the transmitter state-action pair converges or reaches the maximum number of iterations, and the optimal value function of the transmitter state-action pair is obtained Optimal value function decision module, for the transmitter to make decisions according to the current time slot status using the optimal value function of the transmitter state-action pair to make decisions.

8. An electronic device, characterized in that, It includes: A memory that stores at least one instruction; and A processor that executes the instructions stored in the memory to implement the two-dimensional intelligent anti-jamming decision-making method based on the Q-learning algorithm as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the two-dimensional intelligent anti-jamming decision-making method based on the Q-learning algorithm as described in any one of claims 1 to 6.