A communication counter-jamming strategy allocation method based on reinforcement learning

By constructing a wireless communication adversarial scenario model and a deep reinforcement learning model based on reinforcement learning, the allocation of interference resources is optimized, solving the problems of implementation difficulties and low efficiency in traditional methods. This enables fast and reliable allocation of interference strategies and improves interference effectiveness.

CN119997091BActive Publication Date: 2025-10-21XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510154454.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-10-21
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

Traditional jamming strategy allocation methods are difficult to implement in modern communication confrontation, have low allocation efficiency, and are unable to achieve fast and reliable jamming resource allocation under resource-constrained conditions.

Method used

We employ a reinforcement learning-based approach to construct a wireless communication adversarial scenario model, a Markov decision process model, and a deep reinforcement learning model. Through the interaction between the agent and the environment, we optimize the allocation of interference resources, introduce the concept of maximum policy distribution entropy to avoid decisions getting trapped in local optima, and use a twin network to accelerate convergence.

Benefits of technology

It enables rapid and reliable allocation of jamming resources in complex adversarial scenarios, improves jamming effectiveness, reduces decision-making time, and enhances jamming results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119997091B_ABST
    Figure CN119997091B_ABST
Patent Text Reader

Abstract

The application discloses a communication counter-jamming strategy distribution method based on reinforcement learning, and comprises the following steps: S1, establishing a wireless communication counter-jamming scene model; S2, constructing a Markov decision process model interacting with the counter-jamming scene; S3, establishing a jamming resource distribution model comprising an evaluation network, a target network and a strategy network; and S4, combining and training the wireless communication counter-jamming scene model, the Markov decision process model and the deep reinforcement learning model to complete the jamming strategy distribution. The application is based on the reinforcement learning technology, and the reinforcement learning model is constructed and trained, so that the purpose of accurately and quickly realizing the jamming strategy distribution is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of interference strategy allocation, and in particular to a communication anti-interference strategy allocation method based on reinforcement learning. Background Art

[0002] In the context of modern communication confrontation, the application of communication confrontation and counter-countermeasures technologies represented by interference strategy allocation is becoming more and more extensive. As communication confrontation scenarios become more and more complex, higher and higher requirements are placed on the task of interference strategy allocation. However, due to the particularity of confrontation scenarios, it is difficult for the jammer to obtain prior knowledge of the communicator. Traditional interference strategy allocation methods have problems such as implementation difficulties and low allocation efficiency. Therefore, a fast and reliable interference strategy allocation method based on reinforcement learning is needed. Summary of the Invention

[0003] In order to address the shortcomings of existing interference strategy allocation methods, the purpose of the present invention is to provide a communication anti-interference strategy allocation method based on reinforcement learning. Based on reinforcement learning technology, by constructing and training a reinforcement learning model, the purpose of accurately and quickly realizing interference strategy allocation is achieved.

[0004] In order to achieve the above object, the present invention adopts the following technical solutions:

[0005] A communication countermeasure interference strategy allocation method based on reinforcement learning includes the following steps: S1: establishing a wireless communication countermeasure scenario model; S2: constructing a Markov decision process model that interacts with the countermeasure scenario; S3: establishing a deep reinforcement learning model for interference resource allocation that includes an evaluation network, a target network, and a policy network; S4: combining and training the aforementioned wireless communication countermeasure scenario model, the Markov decision process model, and the deep reinforcement learning model to complete the interference strategy allocation.

[0006] In S1, the wireless communication confrontation scenario model is established as follows:

[0007] The wireless communication confrontation scenario model is a many-to-many model, that is, the jammer has multiple jammers, and the communicator has multiple communication links for networking communication, as follows:

[0008] The jammer has N jammers, represented by the set n = {1, 2, ..., N}, and the jammer adopts the aiming jamming mode; the communication party uses the TCP / IP protocol for communication and uses M communication links for network communication. The set of communication links is represented by m = {1, 2, ..., M}. These communication links use equal bandwidth channels that do not interfere with each other and are orthogonal, and the relative importance index of each communication link is represented by W = [ω1, ω2, ..., ω M ];

[0009] In the wireless communication confrontation scenario model, it is assumed that the jammer has mastered the locations of the receivers of each enemy communication link through communication reconnaissance and intelligence analysis, and the model assumes that the positions of each receiver are fixed; for the jammer, the relative importance index of each communication link used by the communicating party is unknown; the jammer hopes to reasonably allocate jamming resources under resource-constrained conditions to obtain greater jamming effectiveness.

[0010] In S2, the constructed Markov decision process model is as follows:

[0011] The reinforcement learning method interacts with the scene and solves the problem by constructing a Markov decision process. The Markov decision process includes the elements of the agent Agent, the state space S, the action space A, the reward function R, and the discount factor γ. The Markov decision process is defined as follows:

[0012] Agent: The jammer formulates a jamming plan through the intelligent jamming engine, which instructs the reconnaissance aircraft to conduct reconnaissance and guides the jammers to conduct coordinated jamming. Therefore, the intelligent jamming engine is considered an agent in the Markov decision process.

[0013] State space S: The environment state S(t) represents the current interference resource allocation scheme and the interference effect of the interference scheme. S(t) is an (N+1)-row M-column matrix consisting of the interference resource allocation matrix X(t) and the interference effect evaluation matrix E(t), namely:

[0014]

[0015] The interference resource allocation matrix is:

[0016] X(t)=[x1(t),x2(t),…,x N (t)] T

[0017] Among them, x i (t)=[c i1 (t) c i2 (t) … c iM (t)], 1≤i≤N represents the jamming target of a single jammer; c ij (t)∈{0,1},c ij (t) = 1 means that the i-th jammer interferes with the j-th communication link, c ij =0 means that the i-th jammer does not interfere with the j-th communication link;

[0018] The interference effect evaluation matrix is:

[0019] E(t)=[τ1(t) τ2(t) … τ M (t)]

[0020] Among them, τ j (t)∈{0,1},τ j (t) = 1 represents the symbol error rate of the jth communication link evaluated by the interference party. Reach the preset value τ0, τ j (t) = 0 means that the symbol error rate does not reach the preset value τ0;

[0021] Action space A: Each jammer selects at most U communication links to interfere at time t, and applies a total power of no more than P on the corresponding channels. max The interference signal is, so the interference strategy of the jammer is:

[0022] A(t)=[a1(t) a2(t)…a N (t)] T

[0023] Among them, a i (t)=[p i1 (t) p i2 (t) … p iM (t)], 1≤i≤N represents the interference resource allocation of the i-th jammer, where 0≤p ij (t)≤P max , 1≤j≤M,p ij (t) = 0 means that the i-th jammer does not interfere with the j-th link, otherwise it means that the i-th jammer interferes with the j-th link and the interference signal power is p ij (t), and satisfies and sign is the sign function.

[0024] Reward function R: The role of the reward function mechanism in reinforcement learning is to tell the agent the relative merits of its current behavior. Therefore, the reward function can guide the optimization direction of the reinforcement learning method. In the interference resource allocation optimization problem of communication confrontation, the jammer's goal is to minimize the interference power while achieving the desired symbol error rate, avoiding excessive power that exposes the jammer's location. The reward function is defined as:

[0025]

[0026] Among them, ω i is the relative importance coefficient of the i-th communication link; sign is the sign function; is the symbol error rate of the i-th link; τ0 is the set symbol error rate threshold; P i (t) is the total interference power on the i-th link;

[0027] The goal of the interference resource allocation optimization problem is to maximize the interference effectiveness of the allocation scheme, that is, to maximize the cumulative reward obtained by the interferer over a period of time in the Markov decision process model:

[0028]

[0029] Where γ∈[0,1] is the discount factor, which indicates the degree of influence of future returns on the current state; T is the time period of the Markov decision process model;

[0030] The training goal of using reinforcement learning to solve the interference resource allocation optimization problem is to find the optimal strategy π * , achieving return G t Expected maximization. That is:

[0031]

[0032] in, Returns the maximum strategy, E[·] means averaging, G t is the reward function.

[0033] In S3, the objective function of the deep reinforcement learning model is constructed by introducing the maximum policy distribution entropy, as follows:

[0034] Deep reinforcement learning is a machine learning method that does not require prior information. It uses a trial-and-error approach to learn. The agent continuously interacts with the environment and takes actions based on the currently learned strategy. These actions change the state of the environment. The agent then updates and corrects the strategy based on the feedback from the environment.

[0035] The policy distribution entropy is a parameter that evaluates the randomness of a policy. A larger policy distribution entropy indicates a stronger randomness of the policy and a stronger ability to explore the environment. Sufficient exploration can prevent the policy from falling into a local optimum. Therefore, the concept of policy distribution entropy is introduced into the target reward function. This maximizes the policy distribution entropy while maximizing the cumulative reward, allowing more strategies to be explored during the policy optimization process.

[0036] The calculation method of strategy distribution entropy is as follows:

[0037] H t =lg(π φ (a t |s t ))

[0038] H t is the current policy distribution entropy, is the current policy;

[0039] By introducing the maximum policy distribution entropy, the objective function of the constructed deep reinforcement model is:

[0040]

[0041] In the formula, argmax represents the strategy of finding the maximum expectation, ρ π is the state-action trajectory distribution formed by the strategy π, s t 、a t , r are the state, action and immediate reward at step t respectively, E represents the mathematical expectation operation, ∑ t r(s t ,a t ) represents the cumulative reward within a certain period, i.e., the cumulative interference effectiveness;

[0042] Recursively solve the optimal strategy π * The Q function iteration formula used is:

[0043] Q(s t ,a t )=r t +γE[Q(s t+1 ,a t+1 )-(1-α)H(t)-αH t+1 ]

[0044] The Q function represents the action value in the current state, and α is the update coefficient of the policy network. By giving an initial value, it can be adaptively updated during the learning process.

[0045] In S3, the established interference resource allocation deep reinforcement learning model including the evaluation network, target network, and policy network is as follows:

[0046] In the deep reinforcement learning model of interference resource allocation, the evaluation network and the policy network are both neural network models, where the policy network is a single network structure, which is used to give the current optimal interference resource allocation plan; the evaluation network and the target network use twin networks, that is, two neural networks with the same structure are used to calculate the value of the allocation plan given by the policy network respectively, and compare the allocation plan values ​​given by the two policy networks. The larger plan value is used to optimize the parameters of the policy network; after training, the evaluation network converges to the optimal value function, and the policy network converges to the optimal strategy, that is, the optimal resource allocation plan is obtained.

[0047] In S3, the evaluation network and target network in the interference resource allocation deep reinforcement learning model including the evaluation network, target network and policy network are as follows:

[0048] The twin evaluation network contains two groups of four networks: the first evaluation network Q1, the target network corresponding to the first evaluation network Q1 The second evaluation network Q2 and the target network corresponding to the second evaluation network Q2

[0049] The update rule for the evaluation network is as follows:

[0050] Step 3.1: Calculate the objective function

[0051]

[0052] in, For the target Q network in s t+1 , The corresponding state action value when For the policy network in state s t+1 Reparameterization is a strategy gradient calculation technique that facilitates derivation. Instead of directly sampling actions from the normal distribution consisting of the mean μ and standard deviation σ of the policy network output, random noise ε that satisfies the normal distribution is introduced into the sampling process, and the action is generated using the formula a = μ + σ·ε.

[0053] Step 3.2: Define the value loss function

[0054]

[0055] Among them, θ is the evaluation network parameter, is the target network parameter, R t is the current state s t Take action a t The instant reward, B is the experience replay batch size;

[0056] Step 3.3: Update the evaluation network parameters using gradient descent

[0057]

[0058] in, is the gradient operator, θ1, θ2 are network parameters, α θ is the learning rate;

[0059] Step 3.4: The target network parameters are updated using a single-step soft update method, where τ is the flexible update coefficient. The update method is as follows:

[0060]

[0061] In S3, the policy network in the interference resource allocation deep reinforcement learning model that includes the evaluation network, target network, and policy network is as follows:

[0062] The policy network uses a single-layer neural network, and the loss function that introduces the maximum policy distribution entropy is as follows:

[0063]

[0064] For the policy network in state s t The reparameterized action outputted below, φ is the policy network parameter, and α is the update coefficient of the policy network;

[0065] The process of updating the policy network parameters using gradient descent is as follows:

[0066]

[0067] α φ is the policy network learning rate, is the gradient operator of the policy network.

[0068] In S4, the aforementioned wireless communication adversarial scenario model, Markov decision process model, and deep reinforcement learning model are combined and trained to complete the jamming strategy allocation. The specific combination and training process is as follows:

[0069] First, the evaluation network, target network, policy network parameters and experience recycling pool are initialized, and then the state space S is initialized. During the training process, the following actions are performed cyclically: the Markov decision process model selects action a according to the state and executes it, obtains the immediate reward value of the current action and the next state, and stores the current action, state, and reward value in the experience recycling pool; when the capacity of the experience recycling pool is greater than the specified number, random sampling is performed from the experience recycling pool, the target value is calculated, and the evaluation network parameters and policy network parameters are updated according to the rules until the preset number of training times is reached.

[0070] Compared with the prior art, the present invention has the following advantages:

[0071] 1. A clear confrontation scenario model was constructed, and the interference resource allocation effect in specific confrontation scenarios was good.

[0072] 2. When constructing the objective function of the deep reinforcement model, the concept of maximum policy distribution entropy is introduced to avoid the decision-making falling into the optimal solution.

[0073] 3. The use of twin evaluation networks can converge to the optimal solution faster and make decisions quickly. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Attachment Figure 1 It is an interference countermeasure model.

[0075] Attachment Figure 2 A framework for reinforcement learning-based interference policy allocation methods. DETAILED DESCRIPTION

[0076] The following describes embodiments of the present application, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below are exemplary and intended to be used to explain the present application, and should not be construed as limiting the present application.

[0077] A communication countermeasure interference strategy allocation method based on reinforcement learning includes the following steps: S1: establishing a wireless communication countermeasure scenario model; S2: constructing a Markov decision process model that interacts with the countermeasure scenario; S3: establishing a deep reinforcement learning model for interference resource allocation that includes an evaluation network, a target network, and a policy network; S4: combining and training the aforementioned wireless communication countermeasure scenario model, the Markov decision process model, and the deep reinforcement learning model to complete the interference strategy allocation.

[0078] In S1, the wireless communication confrontation scenario model is established as follows:

[0079] The wireless communication confrontation scenario model is a many-to-many model, that is, the jammer has multiple jammers, and the communicator has multiple communication links for networking communication, as follows:

[0080] The jammer has N jammers, represented by the set n = {1, 2, ..., N}, and the jammers use a targeted jamming mode. The communicating parties use the TCP / IP protocol for communication and use M communication links for network communication. The set of communication links is represented by m = {1, 2, ..., M}. These communication links use equal bandwidth channels that do not interfere with each other and are orthogonal. The relative importance index of each communication link is represented by W = [ω1, ω2, ..., ω M ].

[0081] In the wireless communication adversarial scenario model, the jammer assumes that, through communication reconnaissance and intelligence analysis, the jammer has determined the locations of the receivers on each of the enemy's communication links. The model assumes that each receiver's location is fixed. For the jammer, the relative importance of each communication link used by the enemy is unknown. Given resource constraints, the jammer aims to rationally allocate jamming resources to maximize jamming effectiveness.

[0082] In the embodiment, the interference countermeasure scenario constructed is as shown in the attached Figure 1 As shown in FIG, the jammer uses a jammer to jam the communication link of the jammed party's receiver.

[0083] In S2, the constructed Markov decision process model is as follows:

[0084] Reinforcement learning methods interact with the scene and solve problems by constructing a Markov decision process. The Markov decision process contains elements such as the agent, the state space S, the action space A, the reward function R, and the discount factor γ. The Markov decision process is defined as follows:

[0085] Agent: The jammer formulates a jamming plan through an intelligent jamming engine, which can guide the reconnaissance aircraft to conduct reconnaissance and guide the jammers to conduct coordinated jamming. Therefore, the intelligent jamming engine can be regarded as an agent in the Markov decision process.

[0086] State space S: The environment state S(t) represents the current interference resource allocation scheme and the interference effect of the interference scheme. S(t) is an (N+1)-row M-column matrix consisting of the interference resource allocation matrix X(t) and the interference effect evaluation matrix E(t), namely:

[0087]

[0088] The interference resource allocation matrix is:

[0089] X(t)=[x1(t),x2(t),…,x N (t)] T

[0090] Among them, x i (t)=[c i1 (t) c i2 (t) … c iM (t)], 1≤i≤N represents the jamming target of a single jammer; c ij (t)∈{0,1},c ij (t) = 1 means that the i-th jammer interferes with the j-th communication link, c ij =0 means that the i-th jammer does not interfere with the j-th communication link.

[0091] The interference effect evaluation matrix is:

[0092] E(t)=[τ1(t) τ2(t) … τ M (t)]

[0093] Among them, τ j (t)∈{0,1},τ j (t) = 1 represents the symbol error rate of the jth communication link evaluated by the interference party. Reach the preset value τ0, τ j (t)=0 indicates that the symbol error rate does not reach the preset value τ0.

[0094] Action space A: Each jammer can choose to interfere with at most U communication links at time t, and apply a total power of no more than P on the corresponding channels. max The interference signal is, so the interference strategy of the jammer is:

[0095] A(t)=[a1(t) a2(t) … a N (t)] T

[0096] Among them, a i (t)=[p i1 (t) p i2 (t) … p iM (t)], 1≤i≤N represents the interference resource allocation of the i-th jammer, where 0≤p ij (t)≤P max , 1≤j≤M. p ij (t) = 0 means that the i-th jammer does not interfere with the j-th link, otherwise it means that the i-th jammer interferes with the j-th link and the interference signal power is p ij (t), and satisfies and sign is the sign function.

[0097] Reward function R: The role of the reward function mechanism in reinforcement learning is to tell the agent the relative merits of its current behavior. Therefore, the reward function can guide the optimization direction of the reinforcement learning method. In the interference resource allocation optimization problem of communication confrontation, the jammer's goal is to minimize the interference power while achieving the desired symbol error rate, avoiding excessive power that would reveal the jammer's location. The reward function is defined as:

[0098]

[0099] Among them, ω i is the relative importance coefficient of the i-th communication link; sign is the sign function; is the symbol error rate of the i-th link; τ0 is the set symbol error rate threshold; P i (t) is the total interference power on the i-th link.

[0100] The goal of the interference resource allocation optimization problem is to maximize the interference effectiveness of the allocation scheme, that is, to maximize the cumulative reward obtained by the interferer over a period of time in the Markov decision process model:

[0101]

[0102] Among them, γ∈[0,1] is the discount factor, which indicates the degree of influence of future returns on the current state; T is the time period of the Markov decision process model.

[0103] The training goal of using reinforcement learning to solve the interference resource allocation optimization problem is to find the optimal strategy π * , achieving return G t Expected maximization. That is:

[0104]

[0105] in, Returns the maximum strategy, E[·] means averaging, G t is the reward function.

[0106] In S1, the objective function of the deep reinforcement learning model is constructed by introducing the maximum policy distribution entropy, as follows:

[0107] Deep reinforcement learning is a machine learning method that does not require prior information. It uses trial and error to learn. The intelligent agent continuously interacts with the environment and takes actions based on the current learned strategy in the environment. The actions taken will change the state of the environment. The intelligent agent then updates and corrects the strategy based on the feedback given by the environment.

[0108] Policy distribution entropy is a parameter that assesses the randomness of a policy. A larger policy distribution entropy indicates a stronger randomness and a stronger ability to explore the environment. Sufficient exploration can help prevent the policy from falling into a local optimum. Therefore, we introduce the concept of policy distribution entropy into the target reward function. This maximizes the policy distribution entropy while maximizing the cumulative reward, allowing more strategies to be explored during the policy optimization process.

[0109] The calculation method of strategy distribution entropy is as follows:

[0110] H t =lg(π φ (a t |s t ))

[0111] H t is the current policy distribution entropy, α is the entropy coefficient, which can be adaptively updated during the learning process by giving an initial value, π φ (a t |s t ) is the current policy.

[0112] By introducing the maximum policy distribution entropy, the objective function of the constructed deep reinforcement model is:

[0113]

[0114] In the formula, argmax represents the strategy of finding the maximum expectation, ρ π is the state-action trajectory distribution formed by the strategy π, st 、a t , r are the state, action and immediate reward at step t respectively, E represents the mathematical expectation operation, ∑ t r(s t ,a t ) represents the cumulative reward within a certain period of time, that is, the cumulative interference effectiveness.

[0115] Recursively solve the optimal strategy π * The Q function iteration formula used is:

[0116] Q(s t ,a t )=r t +γE[Q(s t+1 ,a t+1 )-(1-α)H(t)-αH t+1 ]

[0117] The Q function represents the action value in the current state, and α is the update coefficient of the policy network. By giving an initial value, it can be adaptively updated during the learning process.

[0118] In S3, the established interference resource allocation deep reinforcement learning model including the evaluation network, target network, and policy network is as follows:

[0119] In the deep reinforcement learning model for interference resource allocation, both the evaluation network and the policy network are neural network models. The policy network is a single network structure, used to determine the current optimal interference resource allocation solution. The evaluation network and target network use twin networks, that is, two neural networks with identical structures. The evaluation network and target network each calculate the value of the allocation solution given by the policy network, compare the values ​​of the two policy network allocation solutions, and use the larger value to optimize the policy network parameters. After training, the evaluation network converges to the optimal value function, and the policy network converges to the optimal policy, thus obtaining the optimal resource allocation solution.

[0120] In S3, the evaluation network and target network in the interference resource allocation deep reinforcement learning model including the evaluation network, target network and policy network are as follows:

[0121] The twin evaluation network contains two groups of four networks: the first evaluation network Q1, the target network corresponding to the first evaluation network Q1 The second evaluation network Q2 and the target network corresponding to the second evaluation network Q2

[0122] The update rule for the evaluation network is as follows:

[0123] Step 3.1: Calculate the objective function

[0124]

[0125] in, For the target Q network in s t+1 , The corresponding state action value. For the policy network in state s t+1 Reparameterization is a strategy gradient calculation technique that facilitates differentiation. Instead of directly sampling actions from the normal distribution consisting of the mean μ and standard deviation σ of the policy network output, random noise ε that satisfies the normal distribution is introduced into the sampling, and the action is generated using the formula a = μ + σ·ε.

[0126] Step 3.2: Define the value loss function

[0127]

[0128] Among them, θ is the evaluation network parameter, is the target network parameter, R t is the current state s t Take action a t The instant reward is , and B is the experience replay batch size.

[0129] Step 3.3: Update the evaluation network parameters using gradient descent

[0130]

[0131] in, is the gradient operator, θ1, θ2 are network parameters, α θ is the learning rate.

[0132] Step 3.4: The target network parameters are updated using a single-step soft update method, where τ is the flexible update coefficient. The update method is as follows:

[0133]

[0134] In S3, the policy network in the interference resource allocation deep reinforcement learning model that includes the evaluation network, target network, and policy network is as follows:

[0135] The policy network uses a single-layer neural network, and the loss function that introduces the maximum policy distribution entropy is as follows:

[0136]

[0137] For the policy network in state s t The reparameterized action outputted below, φ is the policy network parameter, and α is the update coefficient of the policy network.

[0138] The process of updating the policy network parameters using gradient descent is as follows:

[0139]

[0140] α φ is the policy network learning rate, is the gradient operator of the policy network (a commonly used mathematical operator).

[0141] In S4, the aforementioned wireless communication adversarial scenario model, Markov decision process model, and deep reinforcement learning model are combined and trained to complete the jamming strategy allocation. The specific combination and training process is as follows:

[0142] First, the evaluation network, target network, policy network parameters, and experience recycling pool are initialized. Then, the state space S is initialized. During training, the following steps are repeated: The Markov decision process model selects action a based on the state and executes it. The immediate reward value and next state for the current action are obtained, and the current action, state, and reward value are stored in the experience recycling pool. When the capacity of the experience recycling pool exceeds the specified amount, random sampling is performed from the experience recycling pool, the target value is calculated, and the evaluation network and policy network parameters are updated according to the rules. This continues until the preset number of training times is reached.

[0143] In the embodiment, the interference strategy allocation method framework based on reinforcement learning is as shown in the attached Figure 2 As shown in the figure, by randomly initializing the evaluation network parameters and the policy network parameters, giving the initial decision, and entering the interference environment for policy verification, the obtained policy and its value are put into the experience recycling pool. The policy and its value in the experience recycling pool are sampled and enter the evaluation network and policy network in turn for parameter update, completing a policy update.

[0144] The training process in this embodiment sets a flexible update coefficient τ = 0.005, the number of interactions per round T = 1000, the experience recycling pool capacity D = 106, the batch sample size B = 256, the discount factor γ = 0.1, and the initial entropy coefficient value α = 1. In this embodiment, after training, the resulting suppression strategy can achieve a nearly 95% interference suppression success rate in a single round.

[0145] The invention provides a communication interference strategy allocation method based on reinforcement learning, which is beneficial to enhancing the interference effect, accelerating the convergence speed of interference strategy allocation, improving the interference level, saving time and cost, and being beneficial to achieving higher-intensity information confrontation. It has important application value for providing interference resource allocation strategies in a short period of time.

[0146] The foregoing description of the present invention is provided to enable any person skilled in the art to implement or use the present invention. Various modifications to the present invention will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the scope of the present invention. Therefore, the present invention is not limited to the examples and designs described herein, but is intended to be embodied in accordance with the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A communication anti-interference strategy allocation method based on reinforcement learning, characterized in that: The method comprises the following steps: S1: establishing a wireless communication adversarial scenario model; S2: constructing a Markov decision process model that interacts with the adversarial scenario; S3: establishing a deep reinforcement learning model for interference resource allocation that includes an evaluation network, a target network, and a policy network; S4: combining and training the aforementioned wireless communication adversarial scenario model, the Markov decision process model, and the deep reinforcement learning model to complete the interference strategy allocation; In S1, the wireless communication confrontation scenario model is established as follows: The wireless communication confrontation scenario model is a many-to-many model, that is, the jammer has multiple jammers, and the communicator has multiple communication links for networking communication, as follows: The interferer has jammers, using a collection Indicates that the jammer uses the aiming jamming mode; the communication party uses the TCP / IP protocol to communicate and uses Communication links are used for networking communication. The set of communication links is expressed as , these communication links use equal bandwidth channels that do not interfere with each other and are orthogonal, and the relative importance index of each communication link is expressed as ; In S3, the objective function of the deep reinforcement learning model is constructed by introducing the maximum policy distribution entropy, as follows: Deep reinforcement learning is a machine learning method that does not require prior information. It uses a trial-and-error approach to learn. The agent continuously interacts with the environment and takes actions based on the currently learned strategy. These actions change the state of the environment. The agent then updates and corrects the strategy based on the feedback from the environment. Introducing the concept of strategy distribution entropy into the target reward function, maximizing the strategy distribution entropy while maximizing the cumulative reward, allowing more strategies to be explored during the strategy optimization process; The calculation method of strategy distribution entropy is as follows: is the current policy distribution entropy, For the current strategy; By introducing the maximum policy distribution entropy, the objective function of the constructed deep reinforcement model is: In the formula, argmax means finding a strategy that maximizes the expectation. For strategy The state-action trajectory distribution formed, 、 、 are the state, action and immediate reward at step t, represents the mathematical expectation operation, It represents the cumulative reward in a certain period of time, that is, the cumulative interference effectiveness; Recursively solve the optimal strategy Used The function iteration formula is: The function represents the value of the action in the current state, is the update coefficient of the policy network, which can be adaptively updated during the learning process by giving an initial value; In S3, the established interference resource allocation deep reinforcement learning model including the evaluation network, target network, and policy network is as follows: In the deep reinforcement learning model for interference resource allocation, both the evaluation network and the policy network are neural network models, where the policy network is a single network structure that is used to provide the current optimal interference resource allocation solution. The evaluation network and the target network use twin networks, that is, two neural networks with the same structure. They calculate the value of the allocation plan given by the policy network respectively, and compare the allocation plan values ​​given by the two policy networks. The larger plan value is used to optimize the parameters of the policy network. After training, the evaluation network converges to the optimal value function, and the policy network converges to the optimal policy, thus obtaining the optimal resource allocation plan. In S3, the evaluation network and target network in the interference resource allocation deep reinforcement learning model including the evaluation network, target network and policy network are as follows: The twin evaluation network contains two groups of four networks: the first evaluation network Q1, the target network corresponding to the first evaluation network Q1 , the second evaluation network Q2 and the target network corresponding to the second evaluation network Q2 ; The update rule for the evaluation network is as follows: Step 3.1: Calculate the objective function in, For the target Q network The corresponding state action value when For the policy network in state Reparameterization action under the policy gradient calculation technique is a convenient way to calculate the policy gradient, that is, not using the mean value of the policy network output and standard deviation The normal distribution formed is directly sampled to obtain the action, but random noise that satisfies the normal distribution is introduced in the adoption , using the formula Produce action; Step 3.2: Define the value loss function in, To evaluate the network parameters, are the target network parameters, Current status Take action Instant rewards, is the experience replay batch size; Step 3.3: Update the evaluation network parameters using gradient descent in, is the gradient operator, are network parameters, is the learning rate; Step 3.4: The target network parameters are updated using a single-step soft update method. is the flexible update coefficient, and the update method is as follows: ; In S3, the policy network in the interference resource allocation deep reinforcement learning model that includes the evaluation network, target network, and policy network is as follows: The policy network uses a single-layer neural network, and the loss function that introduces the maximum policy distribution entropy is as follows: For the policy network in state The reparameterized action of the output is are policy network parameters, is the update coefficient of the policy network; The process of updating the policy network parameters using gradient descent is as follows: is the policy network learning rate, is the gradient operator of the policy network.

2. A communication anti-interference strategy allocation method based on reinforcement learning according to claim 1, characterized in that: In the wireless communication confrontation scenario model, it is assumed that the jammer has mastered the locations of the receivers of each enemy communication link through communication reconnaissance and intelligence analysis, and the model assumes that the positions of each receiver are fixed; for the jammer, the relative importance index of each communication link used by the communicating party is unknown; the jammer hopes to reasonably allocate jamming resources under resource-constrained conditions to obtain greater jamming effectiveness.

3. The communication anti-interference strategy allocation method based on reinforcement learning according to claim 1 is characterized in that: In S2, the constructed Markov decision process model is as follows: Reinforcement learning methods interact with scenarios and solve problems by constructing a Markov decision process; the Markov decision process includes the agent, the state space , action space , reward function and discount factor element; the Markov decision process is defined as follows: Agent: The jammer formulates a jamming plan through the intelligent jamming engine, which instructs the reconnaissance aircraft to conduct reconnaissance and guides the jammers to conduct coordinated jamming. Therefore, the intelligent jamming engine is considered an agent in the Markov decision process. State Space : Environmental status Indicates the allocation scheme of interference resources and the interference effect of the interference scheme at the current moment, The interference resource allocation matrix is and interference effect evaluation matrix Composed of OK Column matrix, that is: The interference resource allocation matrix is: in, , Indicates the jamming target of a single jammer; , Indicates the The jammer is on the Communication links are interfered with, Indicates the The jammer did not Interference with communication links; The interference effect evaluation matrix is: in, , The interference party's evaluation Symbol error rate of communication links Reach the preset value , Indicates that the symbol error rate does not reach the preset value ; Action Space :Each jammer at time At most, selection interference communication links, and apply a total power of no more than The interference signal is, so the interference strategy of the jammer is: in, , Indicates the The interference resource allocation of the jammers, where , , Indicates the The jammer did not interfere with the link, otherwise it means The jammer interferes with links and the interference signal power is , and satisfies , and , , sign is the sign function; Reward Function The role of the reward function mechanism in reinforcement learning is to tell the agent the relative merits of its current behavior. Therefore, the reward function can guide the optimization direction of the reinforcement learning method. In the interference resource allocation optimization problem of communication confrontation, the jammer's goal is to minimize the interference power while achieving the desired symbol error rate, avoiding excessive power that would reveal the jammer's location. The reward function is defined as: in, For the The relative importance coefficient of the communication link; sign is the sign function; For the The symbol error rate of the link; is the symbol error rate threshold value set; For the Total interference power of the links; The goal of the interference resource allocation optimization problem is to maximize the interference effectiveness of the allocation scheme, that is, to maximize the cumulative reward obtained by the interferer over a period of time in the Markov decision process model: in, is the discount factor, which indicates the degree of influence of future benefits on the current state; is the time period in which the Markov decision process model is located; The training objective of using reinforcement learning to solve the interference resource allocation optimization problem is to find the optimal strategy , realize returns Expectation maximization, that is: in, Returns the largest strategy, Indicates the mean value, is the reward function.

4. The communication anti-interference strategy allocation method based on reinforcement learning according to claim 1 is characterized in that: In S4, the aforementioned wireless communication adversarial scenario model, Markov decision process model, and deep reinforcement learning model are combined and trained to complete the jamming strategy allocation. The specific combination and training process is as follows: First, the evaluation network, target network, policy network parameters and experience recovery pool are initialized, and then the state space is initialized. , the following actions are performed cyclically during the training process: the Markov decision process model selects actions based on the state And execute, get the immediate reward value and next state of the current action, store the current action, state, and reward value in the experience recycling pool; when the capacity of the experience recycling pool is greater than the specified number, randomly sample from the experience recycling pool, calculate the target value, and update the evaluation network parameters and policy network parameters according to the rules until the preset number of training times is reached.

Citation Information

Patent Citations

  • Deep reinforcement learning communication interference resource allocation method fused with noise network

    CN115866760A

  • Dynamic interference power distribution method for incomplete perception

    CN118678451A