Communication anti-interference method and system in ad hoc network scene

By adopting the communication anti-interference method of reinforcement learning combined with DDQN algorithm in the self-organized network scenario, the problem of lack of real-time perception and intelligent adaptation in the existing technology is solved, real-time perception and dynamic adjustment of the electromagnetic environment are realized, resource waste is reduced, and communication quality is improved.

CN120017187APending Publication Date: 2025-05-16HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510238745.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing technology lacks real-time perception capabilities in self-organized networking scenarios, which can easily cause waste of resources and lacks intelligent adaptability, and cannot effectively deal with interference in complex electromagnetic environments.

Method used

The communication anti-jamming method using reinforcement learning combined with DDQN algorithm is adopted, and the electromagnetic environment is perceived in real time through dynamic adjustment of transmitters and jammers, communication solutions are optimized, resource waste is reduced, and new interference measures are quickly adapted to.

Benefits of technology

Real-time perception and dynamic adjustment of the electromagnetic environment in the self-organized network scenario is realized, resource waste is reduced, communication quality is improved, and new interference measures can be quickly adapted to complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017187A_ABST
    Figure CN120017187A_ABST
Patent Text Reader

Abstract

The invention provides a communication anti-interference method and system in an ad hoc network scene. The method comprises the following steps: a transmitter detects a channel to obtain channel information; selecting a communication scheme according to the current channel, wherein the communication scheme comprises transmitting power, a transmitting channel and a spatial position; the receiver calculates the bit error rate; the jammer changes an interference scheme according to the bit error rate in combination with reinforcement learning, wherein the interference scheme comprises interference power, an interference channel and a spatial position; and repeating the steps until the bit error rate is greater than a predetermined threshold. According to the communication anti-interference method and system in the ad hoc network scene provided by the invention, the electromagnetic environment can be perceived, the transmitting scheme is dynamically adjusted according to the change of the interference source, the resource waste is reduced as much as possible on the premise of ensuring the communication quality, and the interference source with newly added interference measures can be quickly adapted to ensure the communication quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication technology, and more specifically, relates to a communication anti-interference method and system in a self-organizing network scenario. Background Art

[0002] The transmitter logic of the traditional anti-interference method is: the transmitter changes the transmission plan according to the bit error rate of the previous moment: if the bit error rate of the previous moment meets the communication standard, the transmitter keeps the transmission plan unchanged; if the bit error rate of the previous moment does not meet the communication standard, there is an 80% possibility of changing the communication channel, a 10% possibility of increasing the transmission power by one level, and a 10% possibility of random movement. The traditional anti-interference method has the following disadvantages:

[0003] 1. Lack of real-time perception capability: Traditional technologies mostly work based on fixed modes. For example, frequency hopping technology jumps between multiple frequency bands, and spread spectrum technology spreads the signal over a wider frequency band. These methods rely on predetermined parameters and modes, rather than real-time environmental changes. The system usually cannot dynamically detect and adjust according to changes in the real-time electromagnetic environment and interference sources.

[0004] 2. It is easy to cause waste of resources: Traditional anti-interference technology usually relies on one or several fixed anti-interference measures, such as frequency hopping, spread spectrum or signal encryption. When using anti-interference methods, the system does not achieve the best optimization when dealing with interference, resulting in unnecessary resource consumption, which in turn affects the overall efficiency. For example, frequency hopping requires frequent switching of frequencies, and if the interference is not concentrated in a certain frequency band, frequency hopping becomes a waste of resources;

[0005] 3. Lack of intelligent adaptability: Traditional anti-interference technologies mostly rely on preset algorithms and fixed modes, which means they lack intelligent perception and real-time adaptability to interference environments. Especially in complex electromagnetic environments, fixed anti-interference measures may not be able to adjust their strategies in real time to deal with interference sources with new interference measures, resulting in communication interruption or performance degradation. Summary of the invention

[0006] The purpose of the embodiments of the present invention is to provide a communication anti-interference method in a self-organizing network scenario to solve the technical problems existing in the prior art, such as lack of real-time perception capability, easy waste of resources, and lack of intelligent adaptability.

[0007] To achieve the above object, the technical solution adopted by the present invention is: to provide a communication anti-interference method in a self-organizing network scenario, comprising the following steps:

[0008] The transmitter detects the channel to obtain channel information;

[0009] Selecting a communication scheme according to current channel information, the communication scheme including transmission power, transmission channel and spatial position;

[0010] The receiver calculates the bit error rate;

[0011] The jammer changes the jamming scheme according to the bit error rate combined with reinforcement learning, wherein the jamming scheme includes jamming power, jamming channel and spatial position;

[0012] Repeat the above steps until the bit error rate is greater than a predetermined threshold.

[0013] Optionally, in the step of the transmitter detecting the channel to obtain channel information, the channel information includes the transmission power of the transmitter during the last communication, the channel selected by the transmitter during the last communication, the position of the transmitter during the last communication, and the channel selected by the jammer during this communication. The transmission power of the transmitter, the channel of the transmitter, the position of the transmitter and the channel of the jammer constitute the state of the agent in the environment in the reinforcement learning.

[0014] Optionally, the step of the jammer changing the jamming scheme according to the bit error rate in combination with reinforcement learning includes:

[0015] Preprocessing the state of the agent in the environment using one-hot encoding;

[0016] Using reinforcement learning algorithms to form an anti-interference model based on preprocessed information;

[0017] The interference scheme is changed according to the anti-interference model.

[0018] Optionally, the reinforcement learning algorithm includes a prioritized experience replay algorithm, a DDQN algorithm, and a Dueling DQN algorithm.

[0019] Optionally, the reward and punishment algorithm is set to:

[0020] If the current bit error rate is greater than the previous bit error rate, the reward value is -8;

[0021] If the current bit error rate is greater than 10 -5 When the average bit error rate threshold is exceeded, the reward value is -10, and this round of learning ends and the next round of learning begins.

[0022] Optionally, the training process includes:

[0023] Initialize the environment to form an initialization state;

[0024] Sending the initialization state as the state at this moment into the neural network to obtain the action to be selected by the transmitter;

[0025] The receiver calculates the bit error rate, obtains a reward and penalty algorithm based on the bit error rate, and determines the behavior of the jammer at the next moment based on the bit error rate. The state at this moment, the action to be selected by the transmitter, the reward and penalty value, and the state at the next moment form a quaternion;

[0026] Repeat the above steps until the bit error rate exceeds the predetermined threshold or the communication duration exceeds the predetermined threshold.

[0027] Optionally, after the training process, the testing process includes:

[0028] Loading the trained neural network parameters into the transmitter, and then initializing the environment, initializing the transmitter's transmission power, transmission channel, and the transmitter's position to form an initialization state;

[0029] Send the initialization state as the state at this moment into the neural network to obtain the action to be selected;

[0030] The receiver calculates the bit error rate, obtains a reward and punishment algorithm according to the bit error rate, and determines the behavior of the jammer at the next moment according to the bit error rate;

[0031] Repeat the above steps until the maximum communication duration is reached.

[0032] Optionally, the jammer's jamming scheme is modeled as a Markov model.

[0033] The present invention also provides a communication anti-interference system in a self-organizing network scenario. Using the above-mentioned communication anti-interference method in the self-organizing network scenario, the system includes a communication scene with a predetermined area, a transmitter movable within a predetermined range, a fixed receiver and a jammer. The predetermined range is located within the communication scene, the receiver is located within the predetermined range, and the jammer is movable outside the predetermined range and within the communication scene.

[0034] The beneficial effects of the communication anti-interference method and system in the self-organizing network scenario provided by the present invention are: compared with the prior art, the communication anti-interference method in the self-organizing network scenario of the present invention can implement the perception of the electromagnetic environment, dynamically adjust the transmission plan according to the change of the interference source, and secondly, minimize the waste of resources while ensuring the communication quality. Finally, it can also quickly adapt to the interference source of the newly added interference measures to ensure the communication quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0036] Figure 1 A logical schematic diagram of a jammer provided by an embodiment of the present invention;

[0037] Figure 2 A communication process diagram of a communication anti-interference method in a self-organizing network scenario provided by an embodiment of the present invention;

[0038] Figure 3 An intelligent anti-interference model diagram provided by an embodiment of the present invention;

[0039] Figure 4 A logical diagram of an exploration and utilization algorithm in a self-organizing network scenario provided by an embodiment of the present invention;

[0040] Figure 5 A schematic diagram of an original DQN model provided by an embodiment of the present invention;

[0041] Figure 6 A schematic diagram of a DDQN model provided by an embodiment of the present invention;

[0042] Figure 7 A schematic diagram of the original DQN neural network structure provided by an embodiment of the present invention;

[0043] Figure 8 A schematic diagram of a Dueling_DQN neural network structure is provided for an embodiment of the present invention;

[0044] Fig. 9 Provide a reward logic schematic diagram for an embodiment of the present invention;

[0045] Fig.10 A schematic diagram of a communication anti-interference system in a self-organizing network scenario provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0046] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0047] It should be noted that when an element is referred to as being "fixed to" or "disposed on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element.

[0048] It should be understood that the orientation or position relationship indicated by terms such as "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside" and "outside" are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0049] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0050] The communication anti-interference method in the self-organizing network scenario provided by the embodiment of the present invention is now described.

[0051] Please also read Figures 1 to 3 ,The communication anti-interference method in the self-organizing network scenario includes the following steps:

[0052] The transmitter detects the channel to obtain channel information;

[0053] Selecting a communication scheme according to current channel information, the communication scheme including transmission power, transmission channel and spatial position;

[0054] The receiver calculates the bit error rate;

[0055] The jammer changes the jamming scheme according to the bit error rate combined with reinforcement learning, wherein the jamming scheme includes jamming power, jamming channel and spatial position;

[0056] Repeat the above steps until the bit error rate is greater than a predetermined threshold.

[0057] The communication anti-interference method in the self-organizing network scenario in the above embodiment can implement the perception of the electromagnetic environment, dynamically adjust the transmission plan according to the changes in the interference source, and secondly, minimize the waste of resources while ensuring the communication quality. Finally, it can also quickly adapt to the interference source of the newly added interference measures to ensure the communication quality.

[0058] In some embodiments of the present invention, in the step of the transmitter detecting the channel to obtain channel information, the channel information includes the transmission power of the transmitter during the last communication, the channel selected by the transmitter during the last communication, the position of the transmitter during the last communication, and the channel selected by the jammer during this communication. The transmission power of the transmitter, the channel of the transmitter, the position of the transmitter, and the channel of the jammer constitute the state of the agent in the environment in reinforcement learning.

[0059] The present invention assumes that the transmitter obtains channel information by detecting the channel. After detection, the quantized channel frequency bands are four, of which three frequency bands are available for selection. The transmitter can know the transmitter's transmission power P during the last communication. TRi , the channel C selected by the transmitter during the last communication TRi , the last time you were at the location [x TRi ,y TRi ] and the channel C selected by the jammer for this communication JAMi This constitutes the state[P TR, C TR, C JAM ,x TR ,y TR ], we can see that there are 10*3*7*5*5=5250 possible states.

[0060] See also Figure 1 The jammer can change its jamming strategy according to the bit error rate. The jammer strategy is that if the bit error rate is less than 0.1, there is an 80% probability of changing the jamming channel, a 10% probability of increasing the jamming power, and a 10% probability of moving to the position closest to the transmitter; if the bit error rate is greater than or equal to 0.1, there is a 90% probability of maintaining the original jamming strategy and a 10% probability of reducing the jamming power.

[0061] See also Figure 2 In a communication time slice, the transmitter first detects the channel and obtains the channel C selected by the jammer at this time. JAMi ; Then select the communication scheme at this time according to the current channel situation [P TRi, C TRi, ,x TRi ,y TRi ]; then the bit error rate is calculated at the receiving end; the jammer changes the jamming scheme according to the bit error rate [P JAMi, C JAMi, ,x JAMi ,y JAMi ]. This communication process continues until the bit error rate exceeds the threshold and the communication ends.

[0062] The bit error rate threshold refers to the average bit error rate obtained by performing 1000 communications in each communication time slice, with 1000 randomly generated information bits in each communication. The standard of the average bit error rate threshold is shown in Table 1 below. In order to ensure the communication quality, the threshold is set to 10 -5 .

[0063] Table 1 Average bit error rate threshold table

[0064]

[0065] The present invention assumes that the interference signal has a short change period and its state transition process is a Markov process, and adopts the intelligent anti-interference technology based on DDQN to perform anti-interference.

[0066] See also Figure 3 The present invention uses one-hot encoding to pre-process the state, and then uses the combination of priority experience replay, DDQN and Dueling DQN to upgrade the traditional DQN and build an intelligent anti-interference model.

[0067] In some embodiments of the present invention, the step of the jammer changing the jamming scheme according to the bit error rate combined with reinforcement learning includes:

[0068] The state of the agent in the environment is preprocessed using one-hot encoding;

[0069] Using reinforcement learning algorithms to form an anti-interference model based on preprocessed information;

[0070] The interference scheme is changed according to the anti-interference model.

[0071] In some embodiments of the present invention, the reward and punishment algorithm is set to:

[0072] If the current bit error rate is greater than the previous bit error rate, the reward value is -8;

[0073] If the current bit error rate is greater than 10 -5 When the average bit error rate threshold is exceeded, the reward value is -10, and this round of learning ends and the next round of learning begins.

[0074] Reinforcement learning is a closed-loop system that learns from feedback. The subject of reinforcement learning is the agent. The agent obtains the state by observing the environment and takes action to interact with the environment, thereby obtaining reward information to know whether the action is good or bad. The agent optimizes its strategy based on this reward information and chooses the next action to continue interacting with the environment. The ultimate goal of optimization is to maximize the accumulated reward value.

[0075] In one round, all the rewards of the agent from the beginning to the end of the environment are recorded as R1, R2...R n Define the discount rate as γ∈[0,1]. The definition of the discounted return is

[0076] U t =R t +γR t+1 +γ 2 R t+2 +...+γ n-tR n

[0077] At time t before the round ends, U t is an unknown random variable whose randomness comes from all states and actions after time t.

[0078] The action-value function is defined as:

[0079] Q π (s t ,a t )=E[U t |S t =s t ,A t =a t ]

[0080] This expectation eliminates all states and actions after time t.

[0081] The optimal action-value function is to eliminate the policy π by maximizing:

[0082]

[0083] This means that no matter what strategy π is adopted in the future, the return cannot exceed Q * .Q * The essence of is to be able to foresee the future, and at time t, the expected cumulative reward between time t and time n can be foreseen. Given a state s t , which can make Q * Give each move a score. The move with the highest score will be executed.

[0084] In practice, we get Q * The most effective method is to use a deep Q network DQN, denoted as Q(s,a;ω). The learning goal is that for all s and a, the DQN's prediction Q(s,a;ω) is as close as possible to Q * .

[0085] The most commonly used algorithm for training DQN is the TD (temporal difference) algorithm. t ,a t ,r t ,s t+1 ). For an s t ,predict

[0086] The TD target is as follows

[0087]

[0088] Define the loss function

[0089]

[0090] The gradient of the loss function with respect to ω is

[0091]

[0092] Do another gradient descent update ω:

[0093]

[0094] Here, a is the learning rate, which can be adjusted manually.

[0095] DQN uses a neural network to train to obtain a table of all actions a corresponding to all states s, that is, to eventually make Q converge to Q * .

[0096] Precoding of input data is a key step in deep learning. Its main function is to convert raw data into a form suitable for neural network processing, enhancing the training efficiency, stability and generalization ability of the model. Precoding helps the network effectively capture the key features of the data through digitization, standardization, dimensionality reduction and other means, thereby improving the performance of the model and accelerating the training process.

[0097] If the pre-encoded data has a good numerical distribution (for example, normalized to a similar range), it can help the model converge faster and avoid instability during training. Pre-encoding can ensure the balance of input data and avoid the occurrence of certain features with values ​​that are too large or too small, which affects the learning of the network.

[0098] In deep reinforcement learning (DQN), a scoring table is fitted through a neural network. The input of the neural network is the state State, which needs to be encoded.

[0099] In the present invention, if state[P TR, C TR, C JAM ,x TR ,y TR ] input, for example, state[22.0,[0,0,1,0],[0,1,1,0],-25,-15], after flattening, directly input, there are large values ​​(22.0, -25, -15) and small values ​​(0, 1), the uneven distribution of values ​​will cause the training to not converge for a long time. So it needs to be encoded.

[0100] In this part, the present invention uses one-hot encoding to encode the state. This requires assuming that all states in the state space State are known before one-hot encoding can be used.

[0101] One-hot encoding is a common data preprocessing method, which is mainly used to convert categorical data into numerical data to facilitate processing by machine learning models (especially neural networks). Its basic principle is to convert each category value into a binary vector, in which only the elements of the corresponding category are 1, and the elements in other positions are 0.

[0102] The one-hot encoding process is: the state state has 5 features, namely, transmit power, communication channel, interference channel, x-coordinate and y-coordinate. For the first feature, transmit power, the element of the corresponding category is 1, and the elements of other positions are 0, for example, P TR =2.2, and its corresponding one-hot encoding is [1,0,0,0,0,0,0,0,0,0]. For the second feature, the communication channel, it is directly introduced, for example, C TR =[0,1,0,0], its corresponding one-hot encoding is [0,1,0,0], and so on. Then concatenate the encoding results of all eigenvalues ​​horizontally to get their corresponding one-hot encoding. For example: state[22.0,[0,0,1,0],[0,1,1,0],-25,-15], its corresponding one-hot encoding is [1,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,1,1,0 1,0,0,0,0,0,0,0,0,1].

[0103] Table 2 shows the complete state space.

[0104] Table 2 State space table

[0105]

[0106] In some embodiments of the present invention, the reinforcement learning algorithm includes a prioritized experience replay algorithm, a DDQN algorithm, and a Dueling DQN algorithm.

[0107] In reinforcement learning, exploration and exploitation are very important concepts. Exploration refers to the agent intentionally choosing actions that are not currently optimal during training, but exploring possible environmental states and strategies through some random choices. This process is crucial to the success of reinforcement learning because it helps the agent discover more strategies and environmental states, thereby improving generalization capabilities. Exploitation refers to choosing the currently optimal action.

[0108] The goal of reinforcement learning is to learn a strategy that maximizes long-term cumulative rewards through interaction with the environment. However, if the agent always uses the best action it currently knows, it may fall into a local optimal solution and fail to find the global optimal solution. Exploration is the key means to avoid this situation.

[0109] See also Figure 4 , the ε-greedy algorithm is used in this invention for exploration and utilization. This strategy randomly selects the current optimal action with a probability ε∈[0,1] and randomly selects actions with a probability of 1-ε for exploration. Figure 4 , which is an exploration and utilization algorithm containing a reasoning module.

[0110] Prioritized experience replay is to perform non-uniform random sampling of valuable experience, that is, valuable experience whose predicted value deviates seriously from the true value, to accelerate the training process. When the sample capacity is relatively large, sorting consumes a lot of computing power. Therefore, a binary tree structure is used to store and manage the priority p, which is called the priority tree. The original experience replay pool is still used to store experience.

[0111] Priority experience playback is divided into three parts: the first part is calculation priority; the second part is update priority; and the third part is usage priority.

[0112] (1) Calculation priority

[0113] In the original DQN, experience has serial correlation, and the collected experience cannot be reused, so the experience replay strategy is used. The experience replay strategy randomly and uniformly samples the experience in the experience pool, and cannot perform non-uniform random sampling on valuable experience, that is, valuable experience whose predicted value deviates seriously from the true value. Therefore, priority experience replay is used, with TD error as:

[0114] δ j =q j -y j

[0115] To measure the priority of the sample and modify the corresponding learning rate α. Intuitively, if the probability of an experience being sampled is high, then its corresponding learning rate should be relatively small. The general idea is to sort all samples according to the sample priority. Sample Priority

[0116] p j =|δ j | γ ,γ∈[0,1]

[0117] (2) Update priority

[0118] This function is implemented by the Add() method and the Update() method. The Add() method is used to add new memory samples and

[0119] Priority, and update the priority of the entire tree.

[0120] The pseudo code of the Add() method is as follows.

[0121] Table 3 Add() method pseudo code

[0122]

[0123] The Update() method is used to update the priority of the specified node and recursively update the priorities of all parent nodes.

[0124] The pseudo code of the Update() method is as follows.

[0125] Table 4 Update method pseudo code

[0126]

[0127] (3) Usage priority

[0128] The probability of winning is

[0129]

[0130] Importance sampling weights

[0131] ISWeight j =(b*prob j ) -β ,β∈(0,1)

[0132] Normalized, we get

[0133]

[0134] It can be further deduced that

[0135]

[0136] So, the gradient descent formula becomes

[0137]

[0138] Then sample according to the priority. The sampling method mainly consists of the sample() method and the get_leaf() method. The sample() method is implemented. The get_leaf() method is implemented.

[0139] The sample() method is used to implement this. The pseudo code of the sample() method is shown below.

[0140] Table 5 Pseudo code of the Sample() method

[0141]

[0142] The get_leaf() method is used to compare the random value v with the priority of the node, search from the top of the tree downward until a leaf node is found, and return the index, priority, and stored data of the selected leaf node. The pseudo code of the get_leaf() method is as follows

[0143] Table 6 Pseudo code of the Get_leaf() method

[0144]

[0145] See also Figure 5 , the DDQN algorithm is: The traditional DQN algorithm may over-estimate the value of actions because it uses the same Q network when selecting the optimal action and calculating the target Q value. DDQN uses two networks, one for selecting actions and the other for calculating the target Q value. This effectively reduces this deviation. Specifically, when calculating the target Q value, DDQN uses one Q network to select actions and another Q network to estimate the Q value of the corresponding action, thereby reducing the impact of over-estimation.

[0146] Original DQN, for an s t , predict q t =Q(s t ,a t ;ω), TD target error: Defining the loss function Gradient of the loss function like Underestimate (or overestimate), then y t Underestimation (or overestimation) Underestimate or overestimate.

[0147] ,DDQN algorithm is as follows:DDQN algorithm contains two networks, Evaluation Network: used for real-time learning and updating weights. Target Network: used to generate stable target values, whose weights are regularly copied from the evaluation network and remain unchanged for a period of time. Let the evaluation network eval_q=Q(s,a;ω), the current network parameters are ω and the target network target_q=Q(s,a;ω - )The current network parameters are ω - .

[0148] First do forward propagation on eval_q to get q t =Q(s t ,a t ;ω), then choose Then do forward propagation in target_q and get q t+1 =Q(st+1 ,a * ;ω - ), and then calculate the TD target and TD error, y t =r t +γq t+1 , Back propagate eval_q to get the gradient Do gradient descent to update the parameters of eval_q:

[0149] See also Figure 6 , which is the diagram of DDQN. Assume that Eval_Q(s t+1 ,a1) is the selected a*, then Eval_Q(s t ,a1) is the α to be updated, Target_Q(s t+1 ,a1) is q t+1 .

[0150] Advantages of Dueling DQN algorithm: Dueling DQN decomposes the Q value into state value (V(s)) and advantage function (A(s,a)), so that the algorithm can learn the value of each state more effectively without having to rely solely on the value of each action. This decomposition helps the model focus on the importance of different states more quickly without having to repeatedly evaluate the impact of each action, thereby speeding up the learning process.

[0151] Dueling DQN algorithm is as follows. The original DQN is to optimize the value function Q * is an approximation of the optimal value function Q * The difference between the two lies in the structure of the neural network. It converts the optimal action value Q * Decomposed into the optimal state value V plus the action advantage D. The duel network can learn which states are valuable or worthless without knowing the impact of each action on each state, and is used to handle situations with more actions. In practice, the duel network has better results.

[0152] Through mathematical derivation, we can get the following theorem Q * (s,a)=V * (s)+D * (s,a)-maxD * (s, a), after replacing the neural network, we can get Q(s, a; ω) = V(s; ω V )+D(s,a;ω D )-max(D(s,a;ω D )). In practice, the duel network is defined as follows

[0153] Q(s,a;ω)=V(s;ω V )+D(s,a;ω D )-mean(D(s,a;ω D ))

[0154] See also Figure 7 and Figure 8 , Figure 7 is the structure of the original DQN neural network, Figure 8 This is the structure of the Dueling_DQN neural network.

[0155] Reward is set as follows. Generally speaking, our goal is to make the bit error rate ber as small as possible and the transmission power P TR The smaller the better, the fewer times the channel is switched, the better, the smaller the distance d the transmitter moves, the better, and channel switching is encouraged, while increasing the transmission probability and frequent movement are discouraged. In order to make the different variables have the same order of magnitude, the different variables are normalized.

[0156] Normalized transmit power

[0157]

[0158] Normalized transmitter travel distance

[0159]

[0160] The switching channel loss is artificially set to cost = 0.2.

[0161] When the bit error rate is between 0 and 10 -5 If the current bit error rate is smaller than the previous bit error rate,

[0162] Reward=-((ber)+(norm_P TR )+cost+norm_d)

[0163] This setting has another benefit, which is to encourage the transmitter to switch channels, encourage the transmitter to use lower power communication of [2.2w, 4.4w], and encourage the transmitter to move short distances of 0m or 5m. This is because only when the transmission power or moving distance is in the above state, the negative impact on the reward is lower than the impact of switching channels.

[0164] See also Fig. 9, calculate the reward. If the current bit error rate is greater than the previous bit error rate, then the reward is -8; when the bit error rate is greater than 10^-5, it exceeds the average bit error rate threshold, the reward is -10, and this episode ends and the next episode begins. The advantage of setting the reward in this way is that the maximum reward is 0. When the cumulative reward change rate slows down, it proves that the training is moving in the right direction. When it remains unchanged for a long time, the training is successful.

[0165] In some embodiments of the present invention, the training process includes:

[0166] Initialize the environment to form an initialization state;

[0167] Send the initialization state as the state at this moment into the neural network to obtain the action that the transmitter wants to choose;

[0168] The receiver calculates the bit error rate, obtains the reward and punishment algorithm based on the bit error rate, and determines the behavior of the jammer at the next moment based on the bit error rate. The state at this moment, the action to be selected by the transmitter, the reward and punishment value, and the state at the next moment form a quaternion;

[0169] Repeat the above steps until the bit error rate exceeds the predetermined threshold or the communication duration exceeds the predetermined threshold.

[0170] In some embodiments of the present invention, after the training process, the testing process includes:

[0171] Loading the trained neural network parameters into the transmitter, and then initializing the environment, initializing the transmitter's transmission power, transmission channel, and the transmitter's position to form an initialization state;

[0172] Send the initialization state as the state at this moment into the neural network to obtain the action to be selected;

[0173] The receiver calculates the bit error rate, obtains a reward and punishment algorithm according to the bit error rate, and determines the behavior of the jammer at the next moment according to the bit error rate;

[0174] Repeat the above steps until the maximum communication duration is reached.

[0175] In some embodiments of the present invention, the jammer's jamming scheme is modeled as a Markov model.

[0176] The present invention takes the first time slot of communication as the starting point of an episode in reinforcement learning, and takes the error rate of a certain time slot of communication exceeding 10 -5 The end communication is the end point of an episode in reinforcement learning, and each time slot corresponds to the training step time step in reinforcement learning.

[0177] During training, in each episode, the environment is initialized first, that is, the jammer's interference power, interference channel, and jammer position are initialized: [P JAMi, C JAMi, ,x JAMi ,y JAMi ]; Initialize the transmitter's transmission power, transmission channel, and transmitter location: [P TRi, C TRi, ,x TRi ,y TRi ]

[0178] This constitutes the initial state0[P TR0 ,C TR0 ,C JAM0 ,x TR0 ,y TR0 ]. Send the initial state0 to the neural network q_eval() to get the action0 to be selected [P TR0 ,C TR0 ,x TR0 ,y TR0 ], that is, the action that the transmitter should choose at this moment. The bit error rate ber0 is calculated at the receiving end, and reward0 is obtained from the bit error rate. The bit error rate determines the behavior of the interferer at the next moment [P JAM1 ,C JAM1 ,x JAM1 ,y JAM1 This constitutes the next moment state1[P TR1 ,C TR1 ,C JAM1 ,x TR1 ,y TR1 ]. A four-tuple transition0 is composed of [state0, action0, reward0, state1]. This forms an experience. This is a timestep. The cycle continues until the communication is interrupted or the duration of the timestep exceeds the threshold. This completes an episode. Continue to perform episodes until the training stops. In order to prevent the neural network from overfitting and unnecessary resource consumption due to long training. Here, an early stopping strategy is introduced: take the average duration of the last 5 episodes. If the average is greater than 70, stop training. The following table 7 shows the hyperparameters during training.

[0179] Table 7 Parameter settings during training

[0180]

[0181] The test process involves 600 time slots of communication between the trained transmitter and the traditional anti-jamming transmitter, and the results are compared.

[0182] The logic of the traditional anti-interference transmitter is as follows: the transmitter changes the transmission plan according to the bit error rate of the previous moment: if the bit error rate of the previous moment meets the communication standard, the transmitter keeps the transmission plan unchanged; if the bit error rate of the previous moment does not meet the communication standard, there is an 80% possibility of changing the communication channel, a 10% possibility of increasing the transmission power by one level, and a 10% possibility of random movement.

[0183] In reality, we may encounter the following situation: the transmission scheme of the transmitter is changed on the basis of retaining the original transmission scheme, and the interference scheme of the jammer is also changed on the basis of retaining the original interference scheme, which leads to the situation that the self-organizing network scenario is also changed on the basis of retaining the original one. The new self-organizing network scenario includes not only the original self-organizing network scenario, but also the newly added self-organizing network scenario. This requires that the transmitter that has been trained to adapt to the old scenario can adapt to the new self-organizing network scenario in a shorter time.

[0184] The present invention assumes a new ad hoc network scenario: channels are increased from 3 to 4.

[0185] Table 8 below is a table of communication scenario variables definition for the communication parties in the new environment.

[0186] Table 8 Communication scenario variable definition table for communication parties

[0187]

[0188] Table 9 below is a new scenario interference party communication scenario variable definition table.

[0189] Table 9 Interference party communication scenario variable definition table

[0190]

[0191] Table 10 below is the new scene state space.

[0192] Table 10 State space table

[0193]

[0194] Transfer learning: Transfer learning is a machine learning technique whose core idea is to transfer knowledge learned from one field or task to another field or task. This method is particularly suitable for the following two situations:

[0195] Insufficient data: In the target task, there may be little training data, but there may be a large amount of relevant data in the source task. Transfer learning can use the knowledge in the source task to help the learning of the target task, reducing the dependence on large amounts of data.

[0196] Similarity between tasks: There may be a certain degree of similarity between the original task and the target task. Therefore, the model parameters, feature representations, etc. learned in the original task can be partially applied to the target task, thereby improving the learning efficiency and performance of the target task.

[0197] Anti-interference of new scenes belongs to the category of "similarity between tasks".

[0198] There are several common methods for transfer learning:

[0199] Fine-tuning: This is the most common transfer learning method, often used in deep learning. For example, when training a neural network, it is first pre-trained on a large-scale dataset, then the parameters of the network are transferred to the target task, and then the network is fine-tuned.

[0200] Feature Extraction: In the target task, the feature extractor learned from the source task is used directly without modification. Only new classification or regression layers are added to the target task.

[0201] Zero-shot learning: This method transforms the knowledge of the task into a general representation, which enables prediction without any target task data. It is common in the field of natural language processing.

[0202] The new scene anti-interference adopts the method of fine-tuning. By loading the neural network trained in the old scene into the neural network in the new scene, a shorter training time is performed to achieve anti-interference in the new scene.

[0203] In fine-tuning, maximizing the use of previous network connection parameters is the key to improving model performance. By properly utilizing the parameters of the pre-trained model, you can reduce training time, reduce data requirements, and improve the model's performance on new tasks.

[0204] The present invention adopts a layered unfreezing + layered learning rate adjustment method to maximize the use of previous network connection parameters.

[0205] In terms of the corresponding parameter settings for reinforcement learning, the size of the experience replay pool was reduced (from 2000 to 200) to ensure rapid startup; in terms of random exploration, in order to end free exploration early to ensure the quality of real-time communication, ε-greedy = 0.001, and in order to adapt to the new action space more quickly, during the free exploration stage, 20% of the exploration is done on the old action space and 80% on the new action space; the number of update steps required for the new network to replace the old network is reduced from 200 to 50 to ensure rapid updates.

[0206] The neural network structure of the present invention consists of an input layer l1 (n_feature*768), a hidden layer l2 (768*1024) and an output layer (1024*n_output). There is a dropout layer between l1 and l2 to avoid overfitting. n_feature in the input layer is the code length after state encoding, and n_outputs in the output layer dimension is the size of the action space.

[0207] From the old scene to the new scene, the state and the state space change. One-hot encoding can augment the new scene state while ensuring that the encoding method of the old scene state remains unchanged, and the code length is consistent; the action space changes, which causes the output layer dimension to change (from 750 to 1000).

[0208] The principle of layered unfreezing is that the bottom layers of the original model usually learn common features (such as edges, textures), which are useful for most tasks. The parameters of these layers can be frozen to avoid updating during training. The upper layers usually learn task-specific features. These layers can be unfrozen and fine-tuned on new tasks.

[0209] The process of layered thawing is as follows: at the beginning, the input layer l1, hidden layer l2 and output layer are all frozen. When the training time reaches the threshold for thawing the output layer (the 200th time slot), that is, when the experience fills the experience replay pool, the output layer is thawed and training begins. When the training time reaches the threshold for thawing the hidden layer l2 (the 400th time slot), the hidden layer l2 is thawed. When the training time reaches the threshold for thawing the input layer l1 (the 600th time slot), the input layer l1 is thawed.

[0210] The principle of layered learning rate adjustment is that if the same learning rate is used for all layers, it may cause the underlying parameters to be over-updated and lose the common features of the pre-trained model. Through layered learning rate adjustment, a lower learning rate can be set for the bottom layer to retain the common features, and a higher learning rate can be set for the upper layer to quickly adapt to new tasks.

[0211] The layer-wise learning rates are set as follows: output layer lr=0.001, hidden layer l2=lr*0.5, input layer l1=lr*0.1.

[0212] The training process includes training with transfer learning in a new scenario and direct training without transfer learning in a new scenario. The two trainings are conducted for the same length of time and the training results are compared.

[0213] During transfer learning training, since the output layer dimension changes (from 750 to 1000), the output layer dimension is set, and then the neural network parameters trained in the old scene are loaded. The corresponding weights of the newly added neurons are set to 0.

[0214] In each episode, the environment is initialized first, that is, the jammer's interference power, interference channel, and jammer's location are initialized: [P JAMi, C JAMi, ,x JAMi ,y JAMi ]; Initialize the transmitter's transmission power, transmission channel, and transmitter location: [P TRi, C TRi, ,x TR i,y TR i]

[0215] This constitutes the initial state0[P TR0 ,C TR0 ,C JAM0 ,x TR0 ,y TR0 ]. Send the initial state0 to the neural network q_eval() to get the action0 to be selected [P TR0 ,C TR0 ,x TR0 ,y TR0 ], that is, the action that the transmitter should choose at this moment. The bit error rate ber0 is calculated at the receiving end, and reward0 is obtained from the bit error rate. The bit error rate determines the behavior of the interferer at the next moment [P JAM1 ,C JAM1 ,x JAM1 ,y JAM1 This constitutes the next moment state1[P TR1 ,C TR1 ,C JAM1 ,x TR1 ,y TR1]. A four-tuple transition0 is composed of [state0, action0, reward0, state1]. This forms an experience. This is a timestep. The cycle continues until the communication is interrupted or the duration of the timestep exceeds the threshold. This completes an episode. Continue to perform episodes until the training stops. In order to prevent the neural network from overfitting and unnecessary resource consumption due to excessive training time, an early stopping strategy is introduced here: take the average duration of the last five episodes. If the average is greater than 70, stop training. The following table 2-8 shows the hyperparameters during training.

[0216] Table 11 Parameter settings for transfer learning training in new scenarios

[0217]

[0218] Table 12 Parameter settings for direct training without transfer learning in new scenarios

[0219]

[0220] The test process is: after the training is completed, the transmitter trained by transfer learning and the transmitter trained from scratch are used to conduct communication tests for 600 time slots to compare the performance.

[0221] During the test, the maximum communication duration is set to 600 time slots. Load the trained neural network parameters into the transmitter, and then initialize the environment, that is, initialize the jammer's interference power, interference channel, and jammer location: [P JAMi ,C JAMi ,,x JAMi ,y JAMi ]; Initialize the transmitter's transmission power, transmission channel, and transmitter location: [P TRi ,C TRi ,,x TRi ,y TRi This constitutes the initial state0[P TR0 ,C TR0 ,C JAM0 ,x TR0 ,y TR0 ]. Send the initial state0 to the neural network q_eval() to get the action0 to be selected [P TR0 ,C TR0 ,x TR0 ,y TR0 ], that is, the action that the transmitter should choose at this moment. The bit error rate ber0 is calculated at the receiving end, and reward0 is obtained from the bit error rate. The bit error rate determines the behavior of the interferer at the next moment [P JAM1,C JAM1 ,x JAM1 ,y JAM1 This constitutes the next moment state1[P TR1 ,C TR1 ,C JAM1 ,x TR1 ,y TR1 This is one timestep. The cycle continues until the maximum communication duration is reached.

[0222] The method proposed in the present invention can sense the electromagnetic environment in real time and dynamically adjust the transmission scheme according to the change of the interference source. By adopting the DDQN algorithm in reinforcement learning, the transmitter can learn from the accumulated previous communication experience, and finally ensure that the bit error rate of each time slot is within [0,10^-5], thus ensuring the communication quality.

[0223] The reward algorithm proposed in the present invention can make the transmitter as small as possible in terms of bit error rate ber and transmission power P TR The smaller the better, the fewer times the channel is switched, the better, and the smaller the distance d that the transmitter moves, the better. In addition, channel switching is encouraged, and increasing the transmission probability and frequently moving directions are not encouraged to optimize the transmission plan, ultimately avoiding waste of resources.

[0224] The reinforcement learning anti-interference method of fusion transfer learning proposed in the present invention can enable the transmitter to intelligently perceive and adapt to new communication scenarios in real time, and finally adapt quickly to ensure communication quality. In the new scenario, the transmitter first inherits the model trained in the old scenario, accumulates a small amount of new scenario experience and learns from it while communicating, and adjusts its strategy in real time to cope with the new scenario.

[0225] The present invention has the advantage of real-time perception capability. Relying on the reinforcement learning DDQN algorithm model, the transmitter can perceive the electromagnetic environment in real time, accumulate experience and learn from it, and finally realize dynamic adjustment of the transmission scheme according to the change of the interference source in each time slot, ensuring that the bit error rate is in [0, 10 -5 ] to ensure communication quality.

[0226] Specifically, the jammer's jamming strategy is modeled as a Markov model: the jammer changes its jamming strategy with a certain probability based on the bit error rate obtained by the transmitter's transmission plan in the previous time slot. The transmitter makes a decision based on the channel state observed in this time slot and the transmission plan in the previous time slot. The traditional DQN method is used to generate decisions: learning from experience, using the TD algorithm to train the neural network to fit The optimal action value function Q is approximated by *Finally, three methods are used to upgrade the traditional DQN: Prioritized Experience Replay, Dual Q Learning DDQN, and DuelingDQN. The advantages are: Prioritized Experience Replay: Prioritize learning of valuable experience to accelerate convergence. DDQN: Avoid overestimation. DuelingDQN: Divided into action advantage and state value, regardless of the action, it can accumulate network connection parameters related to state value to accelerate convergence.

[0227] The present invention has the advantage of saving resources. By setting the reward algorithm, the transmitter is made to have a bit error rate ber as small as possible, and the transmission power P TR The smaller the better, the fewer times the channel is switched, the better, and the smaller the distance d that the transmitter moves, the better. In addition, channel switching is encouraged, and increasing the transmission probability and frequently moving directions are not encouraged to optimize the transmission plan, ultimately avoiding waste of resources.

[0228] Specifically, only when the bit error rate is between 0 and 10 -5 If the current bit error rate is smaller than the previous time slot bit error rate, Reward = -((ber) + (norm_P TR )+cost+norm_d). The setting has another benefit, which is to encourage the transmitter to switch channels, encourage the transmitter to use lower power communication of [2.2w, 4.4w,], and encourage the transmitter to move short distances of 0m or 5m. This is because only when the transmission power or moving distance is in the above state, the negative impact on Reward is lower than the impact of switching channels.

[0229] This solution has the advantage of intelligent adaptation. Relying on the reinforcement learning anti-interference method integrated with transfer learning, the transmitter can intelligently perceive and adapt to new communication scenarios in real time, and finally adapt quickly to ensure communication quality.

[0230] In a new scenario, the transmitter first inherits the model trained in the old scenario, accumulates a small amount of new scenario experience while communicating and learns from it in a short period of time, thereby adjusting its strategy in real time so that it can cope with the new scenario.

[0231] Specifically, in the new scenario, the transmitter first loads the model trained in the old scenario for fine-tuning, and freezes all layers of the neural network at the same time; learning begins after the experience replay pool is full, first unfreezing the bottom output layer, and then gradually unfreezing the hidden layer and input layer. Different layers have different learning rates, with the output layer having the highest learning rate, the hidden layer second, and the input layer having the lowest learning rate; training is stopped until the early stopping strategy is reached.

[0232] See also Fig.10The present invention also provides a communication anti-interference system in a self-organizing network scenario, using the communication anti-interference method in a self-organizing network scenario in any of the above embodiments. The communication anti-interference system in the self-organizing network scenario includes a communication scenario with a predetermined area, a transmitter movable within a predetermined range, a fixed receiver and a jammer. The predetermined range is located within the communication scenario, the receiver is located within the predetermined range, and the jammer is movable outside the predetermined range and within the communication scenario.

[0233] See also Fig.10 Based on the characteristics of the ad hoc network without a center, consisting of several nodes, and only short-distance communication without relay nodes and dynamic network topology, this scheme is modeled as follows. The communication scenario of the present invention is a square with a side length of 100m in the suburbs. The nodes are a transmitter, a receiver and a jammer.

[0234] The transmitter's communication scheme consists of changing the transmission power, switching channels, and changing the spatial position in three dimensions. The transmitter can move within a square area with a side length of 50m centered on the receiver. The optional x-axis coordinates are x TR =[-25,-15,-5,5,15], the optional coordinate of the y-axis is y TR =[-25,-15,-5,5,15]. The alternative transmission power is a maximum of 22W, with ten power levels of transmission power P TR The transmitter can detect channels. There are four channel frequency bands detected quantitatively, and three alternative channels, namely C TR =[[1,0,0,0],[0,1,0,0],[0,0,1,0]], which means that one of the first three channels is selected for communication. So the total communication scheme is [P TR, C TR, ,x TR ,y TR ]750 species.

[0235] The receiver is located at [0,0] and remains unchanged.

[0236] The jammer's jamming scheme consists of changing the jamming power, switching channels, and changing the spatial position. The jammer's position is movable, and the coordinates of the x-axis that can be moved are x JAM =[-3,7,17,27,37], the movable coordinate of the y-axis is y JAM =[-3,7,17,27,37]. The maximum value of the candidate interference power is 23W, and there are ten power levels of P JAM .

[0237] There are three interference channels and seven alternative interference schemes, as follows.

[0238] C JAM=[[0,0,0,0],[1,0,0,0],[0,1,0,0],[0,0,1,0],[1,1,0,0],[1,0,1,0],[0,1,1,0]], representing the cases of no interference, one channel interference, and two channels interference. The interference signal is a white noise signal. So the total interference scheme [P JAM ,C JAM ,,x JAM ,y JAM ]There are 1,750 species.

[0239] The gain from the transmitter to the receiver is given by

[0240] G TR_RE =PL TR_RE *G ff *G sf (2-1)

[0241] Due to the existence of path loss in wireless communication, the path loss gain from the transmitter to the receiver is as follows:

[0242] PL TR_RE =d TR_RE -PLfactor ,PLfactor=4 (2-2)

[0243] The signal strength changes due to the multipath effect, resulting in rapid fading. Rapid fading gain G ff The simulation is performed by generating random numbers that follow an exponential distribution, with the formula:

[0244] f(x;λ)=λe -λx ,λ=1 (2-3)

[0245] The signal strength changes caused by obstacles cause shadow fading. The gain of shadow fading is G sf The simulation is performed by generating random numbers that follow a log-normal distribution, with the formula:

[0246]

[0247] Among them, the mean μ of the lognormal distribution is 0, and σ is 8dB. This is the gain of shadow fading, which reflects the change in signal strength caused by obstacles. This communication environment is in the suburbs, and the standard deviation is usually between 6dB and 8dB, so σ is selected as 8dB.

[0248] The gain from the jammer to the receiver is as follows:

[0249] G JAM_RE =PL JAM_RE *G ff *G sf (2-5)

[0250] Due to the existence of path loss in wireless communication, the path loss gain from the transmitter to the receiver is as follows:

[0251] PL JAM_RE =d JAM_RE -PLfactor ,PLfactor=4 (2-6)

[0252] The signal strength changes due to the multipath effect, resulting in rapid fading. Rapid fading gain G ff The simulation is performed by generating random numbers that follow an exponential distribution, with the formula:

[0253] f(x;λ)=λe -λx ,λ=1 (2-7)

[0254] The signal strength changes caused by obstacles cause shadow fading. The gain of shadow fading is G sf The simulation is performed by generating random numbers that follow a log-normal distribution, with the formula:

[0255]

[0256] Among them, the mean μ of the lognormal distribution is 0, and σ is 8dB. This is the gain of shadow fading, which reflects the change in signal strength caused by obstacles. This communication environment is in the suburbs, and the standard deviation is usually between 6dB and 8dB, so σ is selected as 8dB.

[0257] Thus, the signal-to-noise ratio at the receiver is as follows,

[0258]

[0259] Table 13 below is a communication scenario variable definition table for the communication parties.

[0260] Table 13 Communication scenario variable definition table for communication parties

[0261]

[0262] Table 14 below is a table of interfering party communication scenario variables definition.

[0263] Table 14 Interference party communication scenario variable definition table

[0264]

[0265] Table 15 below is a table of other variables definition for communication scenarios.

[0266] Table 15 Definition of other variables in communication scenarios

[0267]

[0268] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A communication anti-interference method in a self-organizing network scenario, characterized in that: The following steps are involved: The transmitter detects the channel to obtain channel information; Selecting a communication scheme according to current channel information, the communication scheme including transmission power, transmission channel and spatial position; The receiver calculates the bit error rate; The jammer changes the jamming scheme according to the bit error rate combined with reinforcement learning, wherein the jamming scheme includes jamming power, jamming channel and spatial position; Repeat the above steps until the bit error rate is greater than a predetermined threshold.

2. The communication anti-interference method in the ad hoc network scenario according to claim 1, characterized in that: In the step of the transmitter detecting the channel to obtain channel information, the channel information includes the transmission power of the transmitter during the last communication, the channel selected by the transmitter during the last communication, the position of the transmitter during the last communication, and the channel selected by the jammer during this communication. The transmission power of the transmitter, the channel of the transmitter, the position of the transmitter, and the channel of the jammer constitute the state of the agent in the environment in the reinforcement learning.

3. The communication anti-interference method in the ad hoc network scenario according to claim 2, characterized in that: The steps of the jammer changing the jamming scheme according to the bit error rate combined with reinforcement learning include: The state of the agent in the environment is preprocessed using one-hot encoding; Using reinforcement learning algorithms to form an anti-interference model based on preprocessed information; The interference scheme is changed according to the anti-interference model.

4. The communication anti-interference method in the ad hoc network scenario according to claim 1, characterized in that: The reinforcement learning algorithms include the Prioritized Experience Replay algorithm, the DDQN algorithm and the Dueling DQN algorithm.

5. The communication anti-interference method in the ad hoc network scenario according to claim 1, characterized in that: The reward and punishment algorithm is set as: If the current bit error rate is greater than the previous bit error rate, the reward value is -8; If the current bit error rate is greater than 10 -5 When the average bit error rate threshold is exceeded, the reward value is -10, and this round of learning ends and the next round of learning begins.

6. The communication anti-interference method in the ad hoc network scenario according to claim 1, characterized in that: The training process includes: Initialize the environment to form an initialization state; Sending the initialization state as the state at this moment into the neural network to obtain the action to be selected by the transmitter; The receiver calculates the bit error rate, obtains a reward and penalty algorithm based on the bit error rate, and determines the behavior of the jammer at the next moment based on the bit error rate. The state at this moment, the action to be selected by the transmitter, the reward and penalty value, and the state at the next moment form a quaternion; Repeat the above steps until the bit error rate exceeds the predetermined threshold or the communication duration exceeds the predetermined threshold.

7. The communication anti-interference method in the ad hoc network scenario according to claim 6, characterized in that: After the training process, the testing process includes: Loading the trained neural network parameters into the transmitter, and then initializing the environment, initializing the transmitter's transmission power, transmission channel, and the transmitter's position to form an initialization state; Send the initialization state as the state at this moment into the neural network to obtain the action to be selected; The receiver calculates the bit error rate, obtains a reward and punishment algorithm according to the bit error rate, and determines the behavior of the jammer at the next moment according to the bit error rate; Repeat the above steps until the maximum communication duration is reached.

8. The communication anti-interference method in the ad hoc network scenario according to claim 1, characterized in that: The jammer's jamming scheme is modeled as a Markov model.

9. A communication anti-interference system in a self-organizing network scenario, using the communication anti-interference method in a self-organizing network scenario according to any one of claims 1 to 8, characterized in that: It includes a communication scene with a predetermined area, a transmitter movable within a predetermined range, a fixed receiver and a jammer, wherein the predetermined range is within the communication scene, the receiver is within the predetermined range, and the jammer is movable outside the predetermined range and within the communication scene.