Communication anti-interference method based on deep reinforcement learning and related device
Through the strategy network of deep reinforcement learning training, the problem of traditional communication anti-interference model learning time and low generalization ability in new interference environments is solved, and fast adaptability and efficient communication anti-interference ability is achieved.
Patent Information
- Application Number
- CN202510338627.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-25
AI Technical Summary
Traditional communication anti-interference models have long learning time in new interference environments, low generalization ability, and it is difficult to quickly adapt to changing interference environments.
Using a deep reinforcement learning method, a decision network is built by establishing a communication anti-interference model, a Markov decision-making process and twin delay deterministic strategy gradient algorithm is used to build a decision network, a generalized policy network is trained, and a TD-DPG algorithm is used to perform q value estimation to enhance the adaptability of the agent.
Improves the adaptability and learning speed of agents in new environments, can quickly solve new task requirements, reduces the risk of overestimation during training and maintains computational stability.
Smart Images

Figure CN120378042A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication anti-jamming, and specifically relates to a training method and related device for a communication anti-jamming model based on deep reinforcement learning of elements. Background Art
[0002] With the advent of the era of information-based warfare, communication jamming and anti-jamming in electronic warfare have become crucial components in modern warfare. Information-based warfare not only changes the combat mode of traditional warfare but also greatly increases its complexity. Currently, the main goal of communication anti-jamming is to dynamically adjust and achieve the security of the electromagnetic spectrum and the ability of asymmetric balance according to the specific communication task requirements and the changes in the electromagnetic jamming environment on the battlefield. Future wars will mainly exhibit the characteristics of network centricity, intelligence, and diversification. At the same time, the technology of communication jamming devices is also constantly advancing. Modern jamming devices not only have basic jamming functions but also integrate intelligent technologies such as sensing, analysis, and learning. Facing this intelligent communication jamming, the traditional anti-jamming communication system is obviously difficult to achieve ideal results. Therefore, future anti-jamming communication systems need to have stronger adaptability, flexibility, and autonomous learning ability, and be able to dynamically adjust according to real-time jamming situations to ensure unobstructed communication.
[0003] The patent document with the application publication number CN115103446A discloses a multi-user communication anti-jamming intelligent decision-making method based on deep reinforcement learning. In the multi-user communication anti-jamming scenario, this method uses deep reinforcement learning and adopts a dynamic ε-greedy strategy to improve the learning rate and accelerate the convergence speed of the algorithm. This invention constructs a wireless communication anti-jamming system model for multiple users. Instead of randomly selecting communication frequency bands through frequency hopping technology, it intelligently selects the optimal communication frequency band for each user according to the current spectrum state with the help of the feedback from the base station. The patent document with the application publication number CN118944769A discloses a self-evolving intelligent communication anti-jamming method and system. This method draws on the biological evolution mechanism and can realize the autonomous evolution of communication waveforms according to the complex electromagnetic environment, dynamically adapt to the electromagnetic environment, and ensure the reliability and effectiveness of communication. Neither of the above two patents considers the problem that the adaptability of the model to the new jamming environment is low when the jamming signal changes, which will cause the model to spend a large amount of redundant time to readjust the system when dealing with similar but different jamming environments. Summary of the Invention
[0004] The object of the present invention is to provide a communication anti-jamming method and related device based on deep reinforcement learning for the problems of long learning time and low generalization ability of traditional communication anti-jamming models in a new interference environment, train a policy network with generalization ability, enhance the adaptability of the intelligent agent to solve communication anti-jamming problems, and the intelligent agent equipped with such a policy network can quickly adapt to the new environment and meet the requirement of quickly solving new tasks with the least samples.
[0005] In order to achieve the above object of the invention, the method of the present invention includes the following steps:
[0006] Step 1: Establish a communication anti-jamming model composed of a transmitter, a receiver and a jammer, and set the spectrum information parameters of the transmitter and the jammer;
[0007] Step 2: Apply the Markov decision process to establish a problem model based on the communication anti-jamming model, and describe the problem model by state, action, transition probability, discount factor and reward;
[0008] Step 3: Based on the problem model, use the twin delayed deterministic policy gradient algorithm to construct a decision network;
[0009] Step 4: Use meta-based reinforcement learning to train the decision network to obtain a trained anti-jamming model;
[0010] Step 5: Input interference data into the trained anti-jamming model to obtain an anti-jamming strategy.
[0011] Further,
[0012] Step 1: Establish a communication anti-jamming model composed of a transmitter, a receiver and a jammer, and set the spectrum information parameters of the transmitter and the jammer;
[0013] Step 2: Apply the Markov decision process to establish a problem model based on the communication anti-jamming model, and describe the problem model by state, action, transition probability, discount factor and reward;
[0014] Step 3: Based on the problem model, use the twin delayed deterministic policy gradient algorithm to construct a decision network;
[0015] Step 4: Use meta-based reinforcement learning to train the decision network to obtain a trained anti-jamming model;
[0016] Step 5: Input interference data into the trained anti-jamming model to obtain an anti-jamming strategy.
[0017] Further, in Step 1, the communication anti-jamming model is:
[0018]
[0019] Where: P x is the set of transmitter transmission powers, and p x,i ∈P x is the transmission power used in the i-th time slot; P b is the set of jammer interference powers, and p b,i ∈P b is the interference power in the i-th time slot; F is the set of transmitter transmission powers, V is the total number of channels; Mod is the set of modulation schemes; Cod is the set of coding rates; v is the set of transmitter transmission rates, Y = Y1×Y2, and the transmission rate v i in the i-th time slot satisfies the formula:
[0020] v i = c i log2m i (2)
[0021] Where: m i ∈Mod is the modulation scheme in the i-th time slot; c i ∈Cod is the coding rate in the i-th time slot.
[0022] Furthermore, step 2 includes:
[0023] Step 2-1: Define the state of the i-th time slot as:
[0024] s i = (f x,i , f b,i , p x,i , p b,i , v i ) (3)
[0025] Where: f x,i and f b,i respectively represent the communication channel of the current time slot and the channel where the interference in the current time slot is located;
[0026] Step 2-2: Define the action of the i-th time slot as:
[0027] a i = (f x,i+1 , p x,i+1 , v i+1 ) (4)
[0028] Step 2-3: Denote the state transition probability as P: S×Α×S → [0,1], which represents the probability of selecting action a i ∈S and transitioning to the next state s i ∈A and s i+1 ∈S, assuming the state transition probability is a definite value, and define the discount factor as γ, where 0 < γ < 1;
[0029] Step 2-4: When the user is in state s i and performs action a i , a corresponding reward value r i will be obtained. The reward set is defined as R, and the reward r i at the i-th time slot is defined as:
[0030]
[0031] where: is the channel switching cost incurred when different communication channels are used in the previous and the next time slots; SINR is the signal-to-interference-plus-noise ratio; J P , J f and J v represent the power switching cost, the channel switching cost, and the transmission rate switching cost respectively; T h,i represents the minimum SINR threshold selected for the i-th time slot.
[0032] Furthermore, Step 3 includes:
[0033] Step 3-1: When establishing a problem model using the Markov decision process, that is, after defining the state, action, transition probability, discount factor, and reward, a decision network consisting of an online actor network, an online twin critic network, a target actor network, and a target twin critic network is constructed based on the twin delayed deterministic policy gradient algorithm.
[0034] Step 3-2: Randomly initialize the parameters of the online network and the target network, perform the outer loop, set the number of loops, then initialize the environment, and obtain the state of the environment as the input;
[0035] Step 3-3: Perform the inner loop, set the number of loops. The agent selects an action according to the current actor network ξ θ and the state, adds exploration noise. After performing the action, the agent obtains the reward and the new state from the environment;
[0036] Step 3-4: Store the current state, action, reward, and new state in the experience replay buffer. Update the current state, randomly sample samples from the experience replay buffer, and update the parameters of the online twin critic network by minimizing a specific loss function;
[0037] Step 3-5: When the time step is a multiple of the update period, update the parameters of the online actor network ξ θ using the policy gradient formula, and update the parameters of the target networks and ;
[0038] Step 3-6: Determine whether the inner loop count has been reached. If the loop count has not been reached, execute Step 3-3; if the loop count has been reached, determine whether the outer loop count has been reached. If the loop count has not been reached, execute Step 3-2; if the loop count has been reached, output the trained online network parameters, including the parameters θ of the online actor network ξ θ and the parameters ε1 and ε2 of the online dual critic network ψ 1ε and ψ 2ε .
[0039] Furthermore, in Step 4, meta-based reinforcement learning is used to train the decision-making network. The specific steps are as follows:
[0040] Step 4-1: After forming the decision-making network with the online actor network, online twin critic network, target actor network, and target twin critic network, randomly initialize the online network parameters and target network parameters;
[0041] Step 4-2: Conduct an outer loop and set the loop count. In each outer loop, sample several tasks from the task distribution;
[0042] Step 4-3: Conduct a task processing loop and set the loop count. Initialize each sampled task, obtain its initial state, and set the actor network parameters of the task to the global policy network parameters. Use the algorithm TD-DPG to update the parameters of the task to obtain the updated parameters;
[0043] Step 4-4: Conduct a task interaction loop and set the loop count. Select and execute an action according to the actor network with updated parameters, then obtain the reward and new state from the environment, store the current state, action, reward, and new state in the experience replay buffer corresponding to the task, and update the current state;
[0044] Step 4-5: Determine whether the task interaction loop count has been reached. If the loop count has not been reached, execute Step 4-4; if the loop count has been reached, update the global meta-parameters;
[0045] Step 4-6: Determine whether the task processing loop count has been reached. If the loop count has not been reached, execute Step 4-3; if the loop count has been reached, determine whether the outer loop count has been reached. If the loop count has not been reached, execute Step 4-2; if the loop count has been reached, obtain the online actor network meta-parameters θ.
[0046] Furthermore, in Step 5, the interference data input into the anti-interference model is:
[0047] s b,k =(f b,k ,p b,k ) (6)
[0048] In the formula: f b,k is the transmission channel of the interference signal, and p b,k is the transmission power of the interference signal;
[0049] The corresponding anti-jamming strategy output by the anti-jamming model is:
[0050] a x,k =(f x,k+1 , p x,k+1 , v x,k+1 ) (7)
[0051] In the formula: f x,k+1 is the communication channel of the transmission signal, p x,k+1 is the transmission power of the transmission signal, and v x,k+1 is the transmission rate.
[0052] A communication anti-jamming device based on deep reinforcement learning, comprising:
[0053] A communication anti-jamming model construction module, configured to establish a communication anti-jamming model composed of a transmitter, a receiver, and a jammer, and set the spectrum information parameters of the transmitter and the jammer;
[0054] A problem model construction module, configured to establish a problem model based on the communication anti-jamming model by applying a Markov decision process, and describe the problem model by state, action, transition probability, discount factor, and reward;
[0055] A decision network construction module, configured to construct a decision network based on the problem model by using the twin delayed deterministic policy gradient algorithm;
[0056] A decision network training module, configured to train the decision network by using meta-based reinforcement learning to obtain a trained anti-jamming model;
[0057] An anti-jamming strategy acquisition module, configured to input interference data into the trained anti-jamming model to obtain an anti-jamming strategy.
[0058] A communication anti-jamming system based on deep reinforcement learning, comprising: a computer-readable storage medium and a processor;
[0059] The computer-readable storage medium is used to store executable instructions;
[0060] The processor is configured to read the executable instructions stored in the computer-readable storage medium and execute the communication anti-jamming method based on deep reinforcement learning.
[0061] A communication anti-jamming system based on deep reinforcement learning, comprising: a computer-readable storage medium and a processor;
[0062] The computer-readable storage medium is used to store executable instructions;
[0063] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the communication anti-jamming method based on deep reinforcement learning.
[0064] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0065] 1. The present invention uses the TD-DPG algorithm to estimate the q value using a dual critic network to solve the problem of overestimating the q value caused by a single critic network. The TD-DPG algorithm uses independent critic networks to provide different value estimates, reducing the variance introduced by a single network and enhancing the evaluation of the policy. During the training process, selecting the smaller q value from the two critic networks as the target value can reduce the risk of overestimating the q value without significantly increasing the computational complexity. Using the copies of the actor network and the dual critic network as the target network and the original network as the online network can enhance the update stability;
[0066] 2. The present invention proposes a meta-based deep reinforcement learning method aimed at training a policy network with generalization ability to enhance the adaptability of the agent to solve the communication anti-jamming problem. An agent equipped with such a policy network can quickly adapt to a new environment and meet the requirement of quickly solving new tasks with a minimum number of samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The present invention will be further described below in conjunction with the drawings and embodiments:
[0068] Figure 1 is a flowchart of a communication anti-jamming method based on deep reinforcement learning according to the present invention;
[0069] Figure 2 is a flowchart of updating the parameters of the decision network based on the TD-DPG algorithm according to the present invention;
[0070] Figure 3 is a flowchart of training the parameters of the decision network based on meta-based reinforcement learning according to the present invention;
[0071] Figure 4 is a system time-frequency diagram under periodic scanning interference according to the present invention;
[0072] Figure 5 is a probability matrix diagram based on which dynamic probability interference is generated according to the present invention;
[0073] Figure 6 is a comparison diagram of cumulative rewards of using the TD-DPG algorithm and the Q-Learning algorithm in an interference environment in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0075] The first aspect of the present invention provides a communication anti-jamming method based on deep reinforcement learning, as Figure 1 shown, including the following steps:
[0076] Step 1: Establish a communication anti-jamming model consisting of a transmitter, a receiver, and a jammer, and set the spectrum information parameters of the transmitter and the jammer;
[0077] Step 2: Apply the Markov decision process to establish a problem model based on the communication anti-jamming model, and describe the problem model by state, action, transition probability, discount factor, and reward;
[0078] Step 3: Based on the problem model, use the twin delayed deterministic policy gradient algorithm to construct a decision network;
[0079] Step 4: Use meta-based reinforcement learning to train the decision network to obtain a trained anti-jamming model;
[0080] Step 5: Input the interference data into the trained anti-jamming model to obtain an anti-jamming strategy.
[0081] In step 1, the communication anti-jamming model is:
[0082]
[0083] In the formula: P x is the set of transmitter transmission powers, p x,i ∈P x is the transmission power used in the i-th time slot; P b is the set of jammer interference powers, p b,i ∈P b is the interference power in the i-th time slot; F is the set of transmitter transmission powers, V is the total number of channels; Mod is the set of modulation methods; Cod is the set of coding rates; v is the set of transmitter transmission rates, Y = Y1×Y2, and the transmission rate v i in the i-th time slot satisfies the formula:
[0084] v i = c i log2m i (2)
[0085] where: m i ∈ Mod is the modulation method of the i-th time slot; c i ∈ Cod is the coding rate of the i-th time slot.
[0086] Step 2 specifically includes:
[0087] Step 2-1: Define the state of the i-th time slot as:
[0088] s i = (f x,i , f b,i , p x,i , p b,i , v i ) (3)
[0089] where: f x,i and f b,i respectively represent the communication channel of the current time slot and the channel where the interference of the current time slot is located;
[0090] Step 2-2: Define the action of the i-th time slot as:
[0091] a i = (f x,i+1 , p x,i+1 , v i+1 ) (4)
[0092] Step 2-3: Denote the state transition probability as P: S × Α × S → [0,1], which represents the probability of selecting action a i ∈ S and transferring to the next state s i ∈ A under the given state s i+1 ∈ S. Assume that the state transition probability is a definite value, and define the discount factor as γ, where 0 < γ < 1;
[0093] Step 2-4: When the user executes action a i in state s i , a corresponding reward value r i will be obtained. The reward set is defined as R. Here, define the signal-to-interference-plus-noise ratio of the i-th time slot as:
[0094]
[0095] where: 0 < δ < 1, representing the attenuation factor of the interference power at the receiving end; α is the channel gain; ι 2 is the noise power; That is, if the communication channel is interfered, then φ(f x,i+1 , f b,i+1 ) is 1, otherwise it is 0.
[0096] When SINRi When it is greater than the minimum SINR threshold value selected according to the actual application, it indicates that the communication in the i-th time slot is successful; otherwise, the communication fails. The reward r of the i-th time slot i is defined as:
[0097]
[0098] In the formula: is the channel switching cost generated when different communication channels are used in the previous time slot and the next time slot; SINR is the signal-to-interference-plus-noise ratio; J P 、J f and J v represent the power switching cost, the channel switching cost, and the transmission rate switching cost respectively; T h,i represents the minimum SINR threshold value selected for the i-th time slot.
[0099] Step 3 includes:
[0100] Step 3-1: Establish a problem model using the Markov decision process. That is, after defining the state, action, transition probability, discount factor, and reward, construct a decision network composed of an online actor network, an online twin critic network, a target actor network, and a target twin critic network based on the twin delayed deterministic policy gradient algorithm. Input the online network parameters, including the actor network ξ θ and the two critic networks ψ 1ε and ψ 2ε as well as the corresponding target networks and Randomly initialize the online network parameters θ, ε1, and ε2, and copy them to the target networks to obtain and
[0101] Step 3-2: Perform the outer loop, set the number of loops to E, initialize the environment, and obtain the s of the environment t as the input;
[0102] Step 3-3: Perform the inner loop, set the number of loops to T, and the agent selects an action a θ according to the current actor network ξ t and the state s t , that is:
[0103] a t = ξ θ (a t ) + ω (7)
[0104] In the formula: ω is the exploration noise, which follows a normal distribution with a mean of 0 and a standard deviation of σ.
[0105] Execute the action a tAfter that, the agent obtains the reward r from the environment t and the new state s t+1 ;
[0106] Step 3-4: Store the current state s t , action a t , reward r t and new state s t+1 in the experience replay buffer D. And update the current state to s t+1 . Randomly sample B samples from the experience replay buffer D, and update the parameters of the online double critic network ψ 1ε and ψ 2ε by minimizing a specific loss function. The specific loss function to be minimized is:
[0107]
[0108] where: ε k is the parameter of the k-th online critic network; D is the experience replay buffer; is the q-value estimated by the k-th online critic network ; y t is the target q-value, and the calculation formula is:
[0109]
[0110] where: r t is the environmental reward at time step t; is the q-value at time step t + 1, estimated by the k-th target critic network ; a t+1 is the action at time step t + 1, obtained by the target actor network α θ based on the state s t+1 ;
[0111] Step 3-5: When the time step t is a multiple of the update period d, update the parameters of the online actor network ξ θ using the policy gradient formula. The policy gradient formula is:
[0112]
[0113] where: α θ is the parameter of the online actor network; is the first online critic network.
[0114] Update the parameters of the target networks and using the update formula:
[0115]
[0116] In the formula: represents the parameters of the target actor network represents the parameters of the k-th target critic network;
[0117] Step 3 - 6: Determine whether the number of inner loop iterations reaches T. If the number of loop iterations has not been reached, execute Step 3 - 3; if the number of loop iterations has been reached, determine whether the number of outer loop iterations reaches E. If the number of loop iterations has not been reached, execute Step 3 - 2; if the number of loop iterations has been reached, output the trained online network parameters, including the parameters θ of the online actor network ξ θ and the parameters ε1 and ε2 of the online double critic networks ψ 1ε and ψ 2ε .
[0118] In Step 4, meta-based reinforcement learning is used to train the decision-making network (as Figure 3 shown), and the specific steps are as follows:
[0119] Step 4 - 1: After forming the decision-making network with the online actor network, the online twin critic network, the target actor network, and the target twin critic network, the input is the online network parameters and the task distribution ρ(M). The online network parameters include the actor network ξ θ and the two critic networks ψ 1ε and ψ 2ε and the corresponding target networks and Randomly initialize the online network parameters θ, ε1, and ε2, and copy the online network parameters θ, ε1, and ε2 to the target networks to obtain and
[0120] Step 4 - 2: Conduct an outer loop and set the number of loop iterations to E. In each outer loop, sample M tasks from the task distribution ρ(M);
[0121] Step 4 - 3: Conduct a task processing loop and set the number of loop iterations to M. Initialize each sampled task M m (m ∈ [1, M]), obtain its initial state s t , and set the actor network parameters θ m of this task to the global θ. Use the algorithm TD - DPG to update the parameters of task M m (as Figure 2 shown) to obtain the updated parameters θ' m ;
[0122] Step 4 - 4: Conduct a task interaction loop and set the number of loop iterations to T. According to the updated parameters θ' m of the actor network, select an action at Execute action a t , and then obtain the reward r from the environment t and the new state s t+1 , store (s t , a t , r t , s t+1 ) into the experience replay buffer D' corresponding to task M m Update the current state to s m ; t+1 ;
[0123] Step 4 - 5: Determine whether the number of inner loop iterations reaches T. If the loop iteration number is not reached, execute Step 4 - 4; if the loop iteration number is reached, update the global meta - parameters according to the experience replay buffer D' of task M m using the formula: m where: υ is the learning rate;
[0124]
[0125] is the gradient calculated based on the experience replay buffers of all tasks;
[0126] Step 4 - 6: Determine whether the number of task - processing loop iterations reaches M. If the loop iteration number is not reached, execute Step 4 - 3; if the loop iteration number is reached, determine whether the number of outer loop iterations reaches E. If the loop iteration number is not reached, execute Step 4 - 2; if the loop iteration number is reached, obtain the meta - parameters θ of the online actor network.
[0127] In Step 5, the interference data input to the anti - interference model is:
[0128] s b,k =(f b,k , p b,k ) (13)
[0129] where: f b,k is the transmission channel of the interference signal, p b,k is the transmission power of the interference signal. The corresponding anti - interference strategy output by the anti - interference model is:
[0130] a x,k =(f x,k+1 , p x,k+1 , v x,k+1 ) (14)
[0131] where: f x,k+1 is the communication channel of the transmission signal, p x,k+1 is the transmission power of the transmission signal, v x,k+1 is the transmission rate.
[0132] Embodiment:
[0133] The embodiments of the present invention are specifically described as follows: The system model includes a transmitter, a receiver, and a jammer. The frequency bands for data transmission and interference are 5 non-overlapping 2-MHz channels between 800 MHz and 810 MHz. The acknowledgment frame is independently transmitted through two control links at 750 MHz and 915 MHz. Other parameter settings are as shown in the following table.
[0134] Table 1
[0135]
[0136]
[0137] The jammer is set to two interference modes: periodic scanning interference and dynamic probability interference. In periodic scanning interference, the jammer sends three interference signals, and each interference signal periodically sweeps through all channels within the frequency band. Figure 4 The shown system time-frequency diagram under this interference, where three different shades of color represent different intensities of interference, sweeps through all channels within the frequency band in a cycle of 5 time slots; in dynamic probability interference, the jammer randomly determines interference modes with different periods according to Figure 5 the shown probability matrix.
[0138] To verify the effectiveness of the present invention in communication anti-jamming, the convergence situation of the present invention is compared with that of a communication anti-jamming system using independent Q-Learning (QL). It can be seen from Figure 6 (1) that the TD-DPG algorithm is superior to the Q-Learning algorithm under periodic scanning interference. The TD-DPG algorithm reaches convergence within 120 cycles, while the Q-Learning algorithm reaches convergence at the 300th cycle. Moreover, after convergence, the reward value of the TD-DPG algorithm converges around 45, while the reward value of the Q-Learning algorithm converges around 38. It can be seen from Figure 6 (2) that the TD-DPG algorithm is superior to the Q-Learning algorithm under dynamic probability interference. The TD-DPG algorithm reaches convergence within about 120 cycles, while the Q-Learning algorithm reaches convergence at the 210th cycle. Moreover, after convergence, the reward value of the TD-DPG algorithm converges around 53, while the reward value of the Q-Learning algorithm converges around 42. Generally speaking, the communication anti-jamming system using the TD-DPG algorithm is superior to the communication anti-jamming system using the Q-Learning algorithm.
[0139] On the other hand, the present invention provides a communication anti-jamming device based on deep reinforcement learning, including:
[0140] A communication anti-jamming model construction module, which is used to establish a communication anti-jamming model composed of a transmitter, a receiver and a jammer, and set the spectrum information parameters of the transmitter and the jammer;
[0141] A problem model construction module, which is used to establish a problem model based on the communication anti-jamming model by applying the Markov decision process, and describe the problem model by state, action, transition probability, discount factor and reward;
[0142] A decision network construction module, which is used to construct a decision network based on the problem model by using the twin delayed deterministic policy gradient algorithm;
[0143] A decision network training module, which is used to train the decision network by using meta-based reinforcement learning to obtain a trained anti-jamming model;
[0144] An anti-jamming strategy acquisition module, which is used to input interference data into the trained anti-jamming model to obtain an anti-jamming strategy.
[0145] On the other hand, the present invention provides a communication anti-jamming system based on deep reinforcement learning, including: a computer-readable storage medium and a processor;
[0146] The computer-readable storage medium is used to store executable instructions;
[0147] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the communication anti-jamming method based on deep reinforcement learning described in the first aspect.
[0148] On the other hand, the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the communication anti-jamming method based on deep reinforcement learning described in the first aspect.
[0149] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0150] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a device for implementing the functions specified in one or more of the blocks.
[0151] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a device for implementing the functions specified in one or more of the blocks.
[0152] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a device for implementing the functions specified in one or more of the blocks.
[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A communication anti-jamming method based on deep reinforcement learning, characterized in that, It includes the following steps: Step 1: Establish a communication anti-jamming model consisting of a transmitter, a receiver, and a jammer, and set the spectrum information parameters of the transmitter and the jammer; Step 2: Apply the Markov decision process to establish a problem model based on the communication anti-jamming model, and describe the problem model by state, action, transition probability, discount factor, and reward; Step 3: Based on the problem model, use the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm to construct a decision network; Step 4: Use meta-based reinforcement learning to train the decision network to obtain a trained anti-jamming model; Step 5: Input interference data into the trained anti-jamming model to obtain an anti-jamming strategy.
2. The communication anti-jamming method based on deep reinforcement learning according to claim 1, wherein: In Step 1, the communication anti-jamming model is: Where: P x is the set of transmitter transmission powers, p x,i ∈P x is the transmission power used in the i-th time slot; P b is the set of jammer jamming powers, p b,i ∈P b is the jamming power in the i-th time slot; F is the set of transmitter transmission powers, V is the total number of channels; Mod is the set of modulation methods; Cod is the set of coding rates; v is the set of transmitter transmission rates, Y = Y1×Y2, and the transmission rate v i in the i-th time slot satisfies the formula: v i = c i log2m i (2) Where: m i ∈ Mod is the modulation mode of the i-th time slot; c i ∈ Cod is the coding rate of the i-th time slot.
3. The communication anti-jamming method based on deep reinforcement learning according to claim 1, characterized in that: Step 2 includes: Step 2-1: Define the state of the i-th time slot as: s i =(f x,i ,f b,i ,p x,i ,p b,i ,v i )(3) where: f x,i and f b,i respectively represent the communication channel of the current time slot and the channel where the interference of the current time slot is located; Step 2-2: Define the action of the i-th time slot as: a i = (f x,i+1 , p x,i+1 , v i+1 ) (4) Step 2-3: Denote the state transition probability as \(P: S\times A\times S\rightarrow[0, 1]\), which represents the probability of selecting an action \(a\in A\) given a state \(s\in S\) and then transitioning to the next state \(s'\in S\). Assume that the state transition probability is a deterministic value, and define the discount factor as \(\gamma\), where \(0 < \gamma < 1\). i \(\in A\) and then transitioning to the next state \(s'\) i \(\in S\). i+1 Step 2-4: When the user performs action a in the s i state, a corresponding reward value r i will be obtained. The reward collection is defined as R, and the reward r i at the i-th time slot is defined as: i In the formula: is the channel switching cost generated when different communication channels are used in the previous time slot and the subsequent time slot; SINR is the signal-to-interference-plus-noise ratio; J P 、J f and J v respectively represent the power switching cost, the channel switching cost, and the transmission rate switching cost; T h,i represents the minimum SINR threshold value selected for the i-th time slot.
4. A communication anti-jamming method based on deep reinforcement learning according to claim 1, characterized in that: Step 3 includes: Step 3-1: After applying the Markov decision process to establish a problem model, that is, after defining the state, action, transition probability, discount factor, and reward, construct a decision network consisting of an online actor network, an online twin critic network, a target actor network, and a target twin critic network based on the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm. Step 3-2: Randomly initialize the online network parameters and the target network parameters, perform an outer loop, set the number of loops, then initialize the environment, and obtain the state of the environment as the input; Step 3-3: Conduct an inner loop and set the number of loop iterations. The agent selects an action according to the current actor network ξ θ and the state, adds exploration noise. After executing the action, the agent obtains a reward and a new state from the environment; Step 3-4: Store the current state, action, reward, and new state in the experience replay buffer. And update the current state, randomly sample samples from the experience replay buffer, and update the parameters of the online twin critic network by minimizing a specific loss function; Step 3-5: When the time step is a multiple of the update period, update the parameters of the online actor network ξ using the policy gradient formula, and update the parameters of the target networks θ and and . θ ; and ; Step 3-6: Determine whether the inner loop count has been reached. If the loop count has not been reached, execute Step 3-3; if the loop count has been reached, determine whether the outer loop count has been reached. If the loop count has not been reached, execute Step 3-2; if the loop count has been reached, output the trained online network parameters, including the parameters θ of the online actor network ξ θ and the parameters ε1 and ε2 of the online double critic networks ψ 1ε and ψ 2ε .
5. A communication anti-jamming method based on deep reinforcement learning according to claim 1, characterized in that: In Step 4, use meta-based reinforcement learning to train the decision network. The specific steps are as follows: Step 4-1: After constructing a decision network consisting of an online actor network, an online twin critic network, a target actor network, and a target twin critic network, randomly initialize the online network parameters and the target network parameters; Step 4-2: Perform an outer loop, set the number of loops, and in each outer loop, sample several tasks from the task distribution; Step 4-3: Perform a task processing loop, set the number of loops, initialize each sampled task, obtain its initial state, and set the actor network parameters of the task to the global policy network parameters. Use the TD-DPG algorithm to update the parameters of the task to obtain the updated parameters; Step 4-4: Perform a task interaction loop, set the number of loops, select and execute an action according to the actor network with updated parameters, then obtain the reward and the new state from the environment, store the current state, action, reward, and new state in the experience replay buffer corresponding to the task, and update the current state; Step 4-5: Determine whether the number of task interaction loops has been reached. If the number of loops has not been reached, execute Step 4-4; if the number of loops has been reached, update the global meta-parameters; Step 4-6: Determine whether the task processing loop count has been reached. If the loop count has not been reached, execute Step 4-3; if the loop count has been reached, determine whether the outer loop count has been reached. If the outer loop count has not been reached, execute Step 4-2; if the outer loop count has been reached, obtain the online actor network meta-parameter θ.
6. The communication anti-jamming method based on deep reinforcement learning according to claim 1, characterized in that: In Step 5, the interference data input into the anti-interference model is: s b,k = (f b,k , p b,k ) (6) where: f b,k is the transmission channel of the interference signal, p b,k is the transmission power of the interference signal; The corresponding anti-interference strategy output by the anti-interference model is: a x,k = (f x,k+1 , p x,k+1 , v x,k+1 ) (7) where: f x,k+1 is the communication channel for the transmitted signal, p x,k+1 is the transmit power of the transmitted signal, v x,k+1 is the transmission rate.
7. A communication anti-jamming device based on deep reinforcement learning, characterized in that, Including: A communication anti-interference model construction module, configured to establish a communication anti-interference model composed of a transmitter, a receiver, and a jammer, and set the spectrum information parameters of the transmitter and the jammer; A problem model construction module, configured to establish a problem model based on the communication anti-interference model by applying a Markov decision process, and describe the problem model by state, action, transition probability, discount factor, and reward; A decision network construction module, configured to construct a decision network based on the problem model by using the twin delayed deterministic policy gradient algorithm; A decision network training module, configured to train the decision network by using meta-based reinforcement learning to obtain a trained anti-interference model; An anti-interference strategy acquisition module, configured to input interference data into the trained anti-interference model to obtain an anti-interference strategy.
8. A communication anti-jamming system based on deep reinforcement learning, comprising: A computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is configured to read the executable instructions stored in the computer-readable storage medium and execute the communication anti-interference method based on deep reinforcement learning according to any one of claims 1-6.
9. A communication anti-jamming system based on deep reinforcement learning, comprising: A computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is configured to read the executable instructions stored in the computer-readable storage medium and execute the communication anti-interference method based on deep reinforcement learning according to any one of claims 1-6.
Citation Information
Patent Citations
Multi-user communication anti-interference intelligent decision-making method based on deep reinforcement learning
CN115103446A
Self-evolution intelligent communication anti-interference method and system
CN118944769A