A frequency hopping intelligent anti-interference decision method based on deep deterministic strategy
Through the HDP-DDPG method, combined with deep deterministic strategies and reinforcement learning algorithms, the problem that existing anti-jammer attacks cannot be effectively avoided in dynamic interference environments is solved, and efficient anti-jamming decisions and stable communication quality are achieved.
Patent Information
- Application Number
- CN202211512206.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-11-29
AI Technical Summary
Existing anti-jamming methods based on deep reinforcement learning may not be able to effectively avoid attacks from intelligent jammers in dynamic interference environments, resulting in communication failure.
A frequency hopping intelligent anti-interference decision-making method based on deep determinism strategy is proposed, called HDP-DDPG. Through the mechanism of mixing dual experience pools and periodic update learning rate, the exploration diversity and convergence speed of the algorithm are enhanced, and the composite experience priority calculation method is designed to improve sample utilization efficiency and avoid local optimization.
It achieves the highest signal-to-noise ratio in complex electromagnetic interference environments, improves the stability and efficiency of anti-interference decisions, and avoids the problems of local optimality and slow convergence speed.
Smart Images

Figure CN116073856B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of frequency hopping anti-interference in wireless communications, and in particular relates to a frequency hopping intelligent anti-interference decision method based on a deep deterministic strategy. Background Art
[0002] With the advancement of science and technology, wireless communication technology has developed by leaps and bounds, and its application scope covers all walks of life. However, due to the openness of the transmission medium, wireless networks are vulnerable to interference attacks. With the development of science and technology, there are more and more interference patterns, and the electromagnetic environment is complex and changeable. When facing these unknown dynamic interferences, traditional anti-interference technology may be completely ineffective. Therefore, the study of intelligent and universal frequency hopping anti-interference decision algorithms is of great significance to improving the quality and security of communication systems.
[0003] Anti-interference decision-making is the core of the anti-interference communication system. The essence of the decision-making process is to adaptively find the optimal solution of the anti-interference strategy in the solution space according to the environmental information and channel quality, under certain constraints, and according to the decision criteria. Since the anti-interference decision is made in a dynamic and random electromagnetic environment, it is essentially a sequential decision-making problem, that is, the transmitter needs to continuously adjust the anti-interference strategy and generate the optimal communication parameters according to environmental changes, and further optimize the anti-interference strategy according to the anti-interference effect. The reinforcement learning algorithm, which has been developing rapidly in recent years, is suitable for and good at solving sequential decision-making problems. It learns autonomously through continuous trial and error interaction with the environment and guiding strategy optimization based on environmental feedback and finally finding the optimal strategy. At the same time, it does not require too much prior information and a large amount of pre-provided training data. Therefore, many scholars have applied reinforcement learning algorithms to the field of communication anti-interference.
[0004] However, the existing anti-interference methods based on deep reinforcement learning methods often use neural network learning strategies to avoid interference. Although good anti-interference effects can be achieved at the current moment, the communication user's previous signal waveform and frequency decision information may have been exposed. If the intelligent jammer obtains the transmitter's communication frequency in advance and applies interference, communication failure will occur. Summary of the invention
[0005] Aiming at the limitations of the anti-interference decision of the existing frequency hopping communication system, the present invention proposes a frequency hopping intelligent anti-interference decision method based on deep deterministic strategy, called HDP-DDPG. Specifically, on the one hand, the model is trained by replaying more experiences with high immediate rewards and large time differential errors to make the model prediction more accurate; on the other hand, the update speed of the network parameters is periodically changed by periodically attenuating the learning rate, the exploration speed is rich and varied, and it is easy to jump out of the local optimum. Finally, the HDP-DDPG network is trained to obtain the final decision model.
[0006] The technical solution adopted by the present invention to solve the technical problem includes the following steps:
[0007] Step 1: Establish a dual-frequency hopping communication system model;
[0008] Step 2: Establish an anti-interference decision model for a dual-frequency hopping communication system;
[0009] Step 3: Transformation of optimization problems based on reinforcement learning;
[0010] Step 4: Anti-interference decision of dual-frequency hopping communication system based on HDP-DDPG;
[0011] Step 5: Train the HDP-DDPG network and output the anti-interference decision model.
[0012] The beneficial effects of the present invention are:
[0013] The present invention formulates the intelligent parameter decision problem in complex electromagnetic interference as a Markov decision process to obtain the highest signal to inference plus noise ratio (SINR). In order to solve it using deep reinforcement learning, continuous states, actions and reward forms are designed according to the optimization problem, and a deep deterministic strategy is proposed to deal with continuous space problems.
[0014] In order to improve the problems of slow convergence and unstable convergence of deep deterministic strategies, the present invention proposes a deep deterministic strategy (HDP-DDPG) that mixes dual experience pools with periodic learning rate updates. The algorithm enhances the exploration diversity of the algorithm through a periodically decaying learning rate. At the same time, a composite experience priority calculation method is designed so that the Agent comprehensively considers immediate rewards and time difference error (TD-error) when selecting experience samples, thereby effectively improving the utilization efficiency of experience samples, avoiding falling into local optimality, and accelerating the convergence speed of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 Schematic diagram of the process of the present invention
[0016] Figure 2 Schematic diagram of simulation evaluation of the present invention and the prior art. DETAILED DESCRIPTION
[0017] The implementation steps of the present invention are further described in detail below.
[0018] like Figure 1 As shown, a frequency hopping intelligent anti-interference decision method based on a deep deterministic strategy specifically includes the following steps:
[0019] Step 1: Establish a dual-frequency hopping communication system model, as follows:
[0020] The mathematical model of a conventional frequency hopping signal is expressed as:
[0021]
[0022] Among them, f c is the minimum frequency hopping frequency, ρ l is the frequency control word generated according to the pseudo-random sequence, which is used to control the change of the frequency hopping frequency. l is the minimum frequency hopping interval, g(t) is the length T c The pulse function, T c is the residence time of each hop, T c During this time, the frequency hopping frequency is based on ρ l to determine the value of .
[0023] The "double-variable" frequency hopping communication technology gives the fixed hopping speed and frequency interval in conventional frequency hopping a time-varying feature. Its main feature is that the interval between each hopping frequency is no longer l It is not an integer multiple of , but any value within the specified range, that is, the minimum hopping frequency interval f l With a time variable f c (a l ) is used instead, where a l is the frequency control word generated by a pseudo-random sequence; the frequency hopping rate v is no longer fixed, but changes pseudo-randomly and nonlinearly at multiple hopping speed levels; accordingly, the dwell time T of each hop c It also changes in pseudo-random nonlinearity, that is, using T c (a l ) to replace T c ; Therefore, the "double-change" frequency-hopping signal can be expressed as:
[0024]
[0025] Assume that the hopping speed v∈[V l ,V u ], the frequency interval d∈[D l ,D u ], then in the kth hop, the user uses the hop speed v k The corresponding residence time T c,k and the frequency hopping frequency f c,k As shown in formula (3) and formula (4) respectively;
[0026]
[0027] f c,k =f c,k-1 ±d k ,dk ∈[D l ,D u ] (4)
[0028] Step 2: Establish an anti-interference decision model for the dual-frequency hopping communication system, as follows:
[0029] Consider a scenario where a pair of transceivers communicate using a dual frequency hopping system in a radio environment with J jammers; at the kth hop, jammer j can arbitrarily select a frequency band The power spectrum density is recorded as Under the guidance of the intelligent agent, the communication user selects a frequency f c,k ∈[F l ,F u ] and send a given power The signal is used for communication; U(f) and BW represent the power spectrum density and bandwidth of the baseband signal respectively; the frequency hopping rate v∈[V l ,V u ], the frequency interval d∈[D l ,D u ], the source rate is b tr ;When interference is detected, the sender avoids interference by changing the frequency hopping parameters of the frequency hopping rate and frequency interval to ensure communication quality;
[0030] In order to improve the communication quality, it is necessary to minimize the user's bit error rate (BER) during communication. Within the Δ time, the bit error rate during communication is expressed by equation (5).
[0031]
[0032] Among them, BER k represents the bit error rate of the kth hop; since the bit error rate is inversely proportional to the signal to interference plus noise ratio at each moment, minimizing the bit error rate is equivalent to maximizing the signal to interference plus noise ratio SINR; therefore, the optimization problem can be expressed as:
[0033]
[0034] Among them, constraint (a) gives the calculation method of signal-to-interference-noise ratio in the kth hop, h k is the average channel gain of the kth hop, p tr is the transmitting power, J k is the total interference power of the kth hop, n kis the total noise power of the kth hop; constraint (b) gives the calculation method of the total interference power; constraint (c) gives the calculation method of the total noise power, n(f) is the power spectral density of Gaussian white noise; constraint (d) indicates that the frequency of the kth hop can be determined by the frequency of the k-1 hop and the frequency interval; constraint (e) indicates that the residence time of the kth hop can be determined by the hop rate.
[0035] Apply reinforcement learning to anti-interference decision-making, making full use of its advantages of not requiring too much prior information and a large amount of training data, and using its continuous interactive trial and error learning results to autonomously learn the optimal anti-interference strategy in a complex and unknown interference environment. In the end, we will learn a state from s k To action a k The optimal mapping strategy a k =μ * (s k ), so that the decision-making agent can make continuous parameter decisions according to this strategy in a continuous period of time in the future and obtain the maximum signal-to-interference-noise ratio.
[0036] Step 3: Transform the optimization problem based on reinforcement learning, as follows:
[0037] In order to obtain the optimal anti-interference strategy μ * , define the communication parameter decision space as a continuous space, and use the DDPG deep reinforcement learning algorithm to solve it; first transform the problem into a Markov decision process; in the Markov decision process, the intelligent body will perceive the current system state, implement actions according to the strategy, thereby changing the state of the environment and obtaining rewards; the following will combine the specific system model to design the parameters of the Markov decision process;
[0038] (1) Action and state space: The current hop number and communication frequency of the user are defined as state parameters. The state is represented by a two-dimensional continuous variable s k =[k,f c,k ], define the action as a two-dimensional continuous variable a k =[v k ,d k ]; k hops later, the user is in state s k =[k,f c,k ], take action a k =[v k ,d k ] and then enter the next state s k+1 =[k+1,f c,k+1 ];
[0039] (2) Reward: Under the guidance of the agent, the user will receive an immediate reward for each step after executing the selected action; the optimization goal is to maximize the signal-to-noise ratio of the system, while the goal of the reinforcement learning algorithm is to maximize the long-term cumulative return expectation E(G k ), defining the long-term cumulative return where γ is the discount factor, r k is the immediate reward for k hops, and the immediate reward is defined as follows:
[0040] r k =SINR k (7).
[0041] Step 4: Anti-interference decision of the dual-frequency hopping communication system based on HDP-DDPG, as follows:
[0042] The network model of HDP-DDPG is as follows Figure 1 As shown in Figure 2, DDPG contains four neural networks: two actor networks with the same structure but different parameters, namely the online actor network μ and the target actor network μ'. The network parameter of the online actor network μ is θ μ , the network parameters of the target actor network μ' are θ μ' target actor network; two critic networks with the same structure but different parameters, namely the online critic network Q and the target critic network Q', where the network parameter of the online critic network Q is θ Q , the network parameters of the target critic network Q' are θ Q' .
[0043] 4-1: Double experience playback. For experience sample i, first define the priority rp based on the immediate reward mechanism i and the priority δp based on the TD-error mechanism i :
[0044] rp i =r i +ε (8)
[0045] δp i =|δ i |+ε (9)
[0046] Among them, r i is the immediate reward of the i-th experience sample; ε is a positive constant used to ensure that each experience has a non-zero priority; δ i is the true value function Q(s i ,ai |θ Q ) and the estimated value y i The difference, called TD-error, is defined as:
[0047] δ i =r i +γQ'(s i+1 ,μ'(s i+1 |θ μ' )|θ Q' )-Q(s i ,a i |θ Q ) (10)
[0048] The samples in the sample pool are sorted according to the priority rp i and δp i Arrange from large to small and get rank1 and rank2. Perform compound sorting on the experience and get:
[0049]
[0050] The composite priorities are:
[0051]
[0052] The parameter η indicates the degree of priority used by the algorithm, and its value range is [0,1]. When η=0, it indicates uniform sampling.
[0053] The probability of sampling experience is defined as:
[0054]
[0055] For experiences with high immediate returns and large time differential errors, the composite ranking is at the front and the probability of being sampled is also high, which can better train the neural network to obtain a better anti-interference parameter strategy.
[0056] 4-2: Learning rate of periodic update; the total number of training rounds M is divided into periods of m, and updated in each period according to the decay law shown in equations (14)-(15):
[0057] k'=mod(m,episode) (14)
[0058] α episode =τ k′ α0 (15)
[0059] β episode =τ k′ β0 (16)
[0060] Where τ is the attenuation factor, α0 and β0 are the initial values of α and β respectively; mod(·) is the modulo operation.
[0061] Step 5: Train the HDP-DDPG network and output the anti-interference decision model, as follows:
[0062] 5-1: Randomly initialize the actor network of HDP-DDPG Weight and critic network Weight
[0063] 5-2: Using θ μ' ←θ μ ,θ Q' ←θ Q Update the weights of the target actor network and the target critic network;
[0064] 5-3: Command Get reward k And enter the next state s k+1 ;
[0065] 5-4: Storage Experience (s k ,a k ,r k ,s k+1 ) to the capacity pool;
[0066] 5-5: With p k is the probability to extract N groups of experience (s n ,a n ,r n ,s n+1 ), where p k Calculated by equations (8)-(13);
[0067] 5-6: Calculate the minimum loss L, where y n =r n +γQ'(s n+1 ,μ'(s i+1 |θ μ' )|θ Q' );
[0068] 5-7: Let k'=mod(m,episode),α episode =τ k′ α0, β episode =τ k′ β0;
[0069] 5-8: Update θ Q :
[0070] 5-9: Update θ μ : in
[0071] 5-10: Update θ μ′ and θ Q′ :θ μ′ =λθ μ′ +(1-λ)θ μ′ and θ Q′ =λθ Q′ +(1-λ)θ Q′ ;
[0072] 5-11: Finally get the optimal anti-interference strategy
[0073] The performance of the method of the present invention is evaluated by simulation, and the simulation parameters are set as follows:
[0074] The source power p of the frequency hopping system tr 150mw, communication frequency band [F l ,F u ]=[0,160]MHz, frequency hopping speed v∈[125,1000]hop / s, channel division interval d∈[1,20]MHz, Gaussian white noise power n0 is 10 -7 mw, the PPER-DQN algorithm was selected as the comparison algorithm in the simulation evaluation. To better demonstrate the algorithm training effect, the rewards and averages were obtained every 20 training rounds, and the 2000 iteration convergence curve was equivalent to the 100 iteration convergence curve. The comparison results are shown in Figure 2. Figure 2 shown.
[0075] Depend on Figure 2 It can be seen that in the convergence stage, the SINR of HDP-DDPG can reach about 28dB, while that of PPER-DQN is only about 16dB, and HDP-DDPG converges more stably with smaller fluctuations.
Claims
1. A frequency hopping intelligent anti-interference decision method based on deep deterministic strategy, characterized by The following steps are involved: Step 1: Establish a dual-frequency hopping communication system model; Step 2: Establish an anti-interference decision model for a dual-frequency hopping communication system; Step 3: Optimization problem conversion based on reinforcement learning; Step 4: Anti-interference decision of dual-frequency hopping communication system based on HDP-DDPG; Step 5: Train the HDP-DDPG network and output the anti-interference decision model; The establishment of the dual-frequency hopping communication system model described in step 1 is as follows: The mathematical model of the frequency hopping signal is expressed as: Among them, f c is the minimum frequency hopping frequency, ρ l is the frequency control word generated according to the pseudo-random sequence, which is used to control the change of the frequency hopping frequency. l is the minimum frequency hopping interval, g(t) is the length T c The pulse function, T c is the residence time of each hop, T c During this time, the frequency hopping frequency is based on ρ l to determine the value of; Minimum hopping frequency interval f l With a time variable f c (ρ l ) is used instead, where ρ l is the frequency control word generated by a pseudo-random sequence; the frequency hopping rate v is no longer fixed, but changes pseudo-randomly and nonlinearly at multiple hopping speed levels; accordingly, the dwell time T of each hop c In the pseudo-random nonlinear change, that is, T c (ρ l ) to replace T c ; Therefore, the double-variable frequency hopping signal can be expressed as: Assume that the hopping speed v∈[V l ,V u ], the frequency interval d∈[D l ,D u ], then in the kth hop, the user uses the hop speed v k The corresponding residence time T c,k and the frequency hopping frequency f c,k As shown in formula (3) and formula (4) respectively; f c,k =f c,k-1 ±d k ,d k ∈[D l ,D u ] (4) The establishment of the anti-interference decision model of the dual-frequency hopping communication system described in step 2 is as follows: Consider a scenario where a pair of transmitting and receiving users use a dual-frequency hopping system to communicate in a radio environment with J jammers; at the kth hop, jammer j randomly selects a frequency band The power spectrum density is recorded as Under the guidance of the intelligent agent, the communication user selects a frequency f c,k ∈[F l ,F u ] and send a given power The signal is used for communication; U(f) and BW represent the power spectrum density and bandwidth of the baseband signal respectively; the frequency hopping rate v∈[V l ,V u ], the frequency interval d∈[D l ,D u ], the source rate is b tr ;When interference is detected, the sender avoids interference by changing the frequency hopping parameters of the frequency hopping rate and frequency interval to ensure communication quality; In the Δ time, the bit error rate during the communication process is expressed by formula (5); Among them, BER k represents the bit error rate of the kth hop; since the bit error rate is inversely proportional to the signal to interference plus noise ratio at each moment, minimizing the bit error rate is equivalent to maximizing the signal to interference plus noise ratio SINR; therefore, the optimization problem is expressed as: Among them, constraint (a) gives the calculation method of signal-to-interference-noise ratio in the kth hop, h k is the average channel gain of the kth hop, p tr is the transmitting power, J k is the total interference power of the kth hop, n k is the total noise power of the k-th hop; constraint (b) gives the calculation method of the total interference power; constraint (c) gives the calculation method of the total noise power, n(f) is the power spectrum density of Gaussian white noise; constraint (d) indicates that the frequency of the k-th hop can be determined by the frequency of the k-1 hop and the frequency interval; constraint (e) indicates that the residence time of the k-th hop can be determined by the hop speed; Apply reinforcement learning to anti-interference decision-making, and use the learning results of continuous interactive trial and error to autonomously learn the optimal anti-interference strategy in a complex and unknown interference environment; ultimately, learn a state from state s k To action a k The optimal mapping strategy a k =μ * (s k ), so that the decision-making agent can make continuous parameter decisions according to this strategy in a continuous period of time in the future and obtain the maximum signal-to-interference-noise ratio; The optimization problem transformation based on reinforcement learning described in step 3 is as follows: In order to obtain the optimal anti-interference strategy μ * , define the communication parameter decision space as a continuous space, and use the DDPG deep reinforcement learning algorithm to solve it; first transform the problem into a Markov decision process; in the Markov decision process, the intelligent body will perceive the current system state, implement actions according to the strategy, thereby changing the state of the environment and obtaining rewards; the following will combine the specific system model to design the parameters of the Markov decision process; (1) Action and state space: The current hop number and communication frequency of the user are defined as state parameters. The state is represented by a two-dimensional continuous variable s k =[k,f c,k ], define the action as a two-dimensional continuous variable a k =[v k ,d k ]; k hops later, the user is in state s k =[k,f c,k ], take action a k =[v k ,d k ] and then enter the next state s k+1 =[k+1,f c,k+1 ]; (2) Reward: Under the guidance of the agent, the user will receive an immediate reward for each step after executing the selected action; the optimization goal is to maximize the signal-to-noise ratio of the system, while the goal of the reinforcement learning algorithm is to maximize the long-term cumulative return expectation E(G k ), defining the long-term cumulative return where γ is the discount factor, r k is the immediate reward for k hops, and the immediate reward is defined as follows: r k =SINR k (7) The anti-interference decision of the dual-frequency hopping communication system based on HDP-DDPG described in step 4 is as follows: The network model of HDP-DDPG contains four neural networks: two actor networks with the same structure but different parameters, namely the online actor network μ and the target actor network μ′; the network parameter of the online actor network μ is θ μ , the network parameters of the target actor network μ′ are θ μ′ ; Two critic networks with the same structure but different parameters, namely the online critic network Q and the target critic network Q', where the network parameter of the online critic network Q is θ Q , the network parameters of the target critic network Q′ are θ Q′ ; 4-1: Double experience playback; for experience sample i, first define the priority rp based on the immediate reward mechanism i and the priority δp based on the TD-error mechanism i : rp i =r i +ε (8) δp i =|δ i |+e (9) Among them, r i is the immediate reward of the i-th experience sample; ε is a positive constant used to ensure that each experience has a non-zero priority; δ i is the true value function Q(s i , a i |θ Q ) and the estimated value y i The difference, called TD-error, is defined as: d i =r i +γQ'(s i+1 ,μ′(s i+1 |θ μ′ )|θ Q′ )-Q(s i ,a i |θ Q ) (10) where y i =r i +γQ′(s i+1 ,μ′(s i+1 |θ μ′ )|θ Q′ ), sort the samples in the sample pool according to the priority rp i and δp i Arrange from large to small and get rank1 and rank2, and perform compound sorting on the experience and get: in The composite priorities are: The parameter η indicates the degree of priority used by the algorithm, and its value range is [0, 1]. When η = 0, it indicates uniform sampling. The probability of sampling experience is defined as: 4-2: Learning rate of periodic update; the total number of training rounds M is divided into periods of m, and updated in the episode according to the decay law shown in equations (14)-(15): k′=mod(m,episode) (14) a episode =t k′ a0 (15) b episode =t k′ β0 (16) Where τ is the attenuation factor, α0 and β0 are the initial values of α and β respectively; mod(·) is the modulus operation; The training HDP-DDPG network described in step 5 outputs the anti-interference decision model, which is as follows: 5-1: Randomly initialize the actor network of HDP-DDPG Weight and critic network Weight 5-2: Using θ μ′ ←θ μ ,θ Q′ ←θ Q Update the weights of the target actor network and the target critic network; 5-3: Order Get reward k And enter the next state s k+1 ; 5-4: Storage Experience (s k , a k , r k ,s k+1 ) to the capacity pool; 5-5: With p k is the probability to extract N groups of experience (s n , a n , r n ,s n+1 ), where p k Calculated by equations (8)-(13); 5-6: Calculate the minimum loss L, where y n =r n +γQ′(s n+1 ,μ′(s i+1 |θ μ′ )|θ Q′ ); 5-7: Let k′=mod(m,episode), and episode =t k′ a0, b episode =t k′ β0; 5-8: Update θ Q : 5-9: Update θ μ : in 5-10: update μ′ and Q′ :i μ′ =λθ μ′ +(1-λ)θ μ′ and Q′ =λθ Q′ +(1-λ)θ Q′ ; 5-11: Finally get the optimal anti-interference strategy