Communication pair anti-interference strategy distribution method based on reinforcement learning
By constructing and training the communication adversarial scenario model based on reinforcement learning, the existing interference strategy allocation method is solved, and efficient interference resource allocation and rapid decision-making are achieved.
Patent Information
- Application Number
- CN202510154454.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The existing interference strategy allocation methods have problems of implementation difficulties and low allocation efficiency in modern communication adversarial scenarios, especially when the interference party finds it difficult to obtain prior knowledge of the communication party.
Using a reinforcement learning-based approach, these models are combined and trained to achieve fast and accurate allocation of interference strategies by constructing wireless communication adversarial scenario models, Markov decision-making process models and deep reinforcement learning models.
It realizes efficient interfering resource allocation under resource constraints, improves interference efficiency, avoids decision-making from falling into local optimality, and can quickly converge to the optimal solution.
Smart Images

Figure CN119997091A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of interference strategy allocation, and in particular to a communication anti-interference strategy allocation method based on reinforcement learning. Background Art
[0002] In the context of modern communication confrontation, the application of communication confrontation and anti-counterattack technologies represented by interference strategy allocation is becoming more and more extensive. As the communication confrontation scenarios become more and more complex, higher and higher requirements are placed on the task of interference strategy allocation. However, due to the particularity of the confrontation scenarios, it is difficult for the jammer to obtain the prior knowledge of the communicator. Traditional interference strategy allocation methods have problems such as difficulty in implementation and low allocation efficiency. A fast and reliable interference strategy allocation method based on reinforcement learning is needed. Summary of the invention
[0003] In order to solve the shortcomings of the existing interference strategy allocation method, the purpose of the present invention is to provide a communication anti-interference strategy allocation method based on reinforcement learning, based on reinforcement learning technology, by constructing and training a reinforcement learning model, to achieve the purpose of accurate and rapid interference strategy allocation.
[0004] In order to achieve the above object, the present invention adopts the following technical solution:
[0005] A communication countermeasure interference strategy allocation method based on reinforcement learning comprises the following steps: S1: establishing a wireless communication countermeasure scenario model; S2: constructing a Markov decision process model that interacts with the countermeasure scenario; S3: establishing a deep reinforcement learning model for interference resource allocation including an evaluation network, a target network and a strategy network; S4: combining and training the aforementioned wireless communication countermeasure scenario model, the Markov decision process model and the deep reinforcement learning model to complete the interference strategy allocation.
[0006] In S1, the wireless communication confrontation scenario model is established as follows:
[0007] The wireless communication confrontation scenario model is a many-to-many model, that is, the jammer has multiple jammers, and the communicator has multiple communication links for networking communication, as follows:
[0008] The jammer has N jammers, represented by the set n = {1, 2, ..., N}, and the jammer adopts the aiming jamming mode; the communicator uses the TCP / IP protocol for communication and uses M communication links for networking communication. The set of communication links is represented by m = {1, 2, ..., M}. These communication links use equal bandwidth channels that do not interfere with each other and are orthogonal, and the relative importance index of each communication link is represented by W = [ω1, ω2, ..., ω M ];
[0009] In the wireless communication confrontation scenario model, it is assumed that the jammer has mastered the locations of the receivers of each communication link of the enemy through communication reconnaissance and intelligence analysis, and the model assumes that the locations of each receiver are fixed; for the jammer, the relative importance index of each communication link used by the communicator is unknown; the jammer hopes to reasonably allocate jamming resources under resource-constrained conditions to obtain greater jamming effectiveness.
[0010] In S2, the constructed Markov decision process model is as follows:
[0011] The reinforcement learning method interacts with the scene and solves the problem by constructing a Markov decision process; the Markov decision process includes the agent Agent, the state space S, the action space A, the reward function R and the discount factor γ elements; the Markov decision process is defined as follows:
[0012] Agent: The jammer formulates a jamming plan through the intelligent jamming engine, and the intelligent jamming engine guides the reconnaissance aircraft to conduct reconnaissance and guides the jammers to conduct coordinated jamming. Therefore, the intelligent jamming engine is regarded as an agent in the Markov decision process;
[0013] State space S: The environment state S(t) represents the allocation scheme of interference resources and the interference effect of the interference scheme at the current moment. S(t) is a (N+1)-row M-column matrix composed of the interference resource allocation matrix X(t) and the interference effect evaluation matrix E(t), that is:
[0014]
[0015] The interference resource allocation matrix is:
[0016] X(t)=[x1(t),x2(t),…,x N (t)] T
[0017] Among them, x i (t) = [c i1 (t) c i2 (t) … c iM (t)], 1≤i≤N represents the jamming target of a single jammer; c ij (t)∈{0,1},c ij (t) = 1 means that the i-th jammer interferes with the j-th communication link, c ij =0 means that the i-th jammer does not interfere with the j-th communication link;
[0018] The interference effect evaluation matrix is:
[0019] E(t)=[τ1(t) τ2(t) … τ M (t)]
[0020] Among them, τ j (t)∈{0,1},τ j (t) = 1 represents the symbol error rate of the jth communication link evaluated by the interferer Reach the preset value τ0, τ j (t) = 0 means that the symbol error rate does not reach the preset value τ0;
[0021] Action space A: Each jammer at time t can jam at most U communication links and exert a total power of no more than P on the corresponding channels. max The interference signal is , so the interference strategy of the interferer, that is, the interference action, is:
[0022] A(t)=[a1(t) a2(t)…a N (t)] T
[0023] Among them, a i (t) = [p i1 (t) p i2 (t) … p iM (t)], 1≤i≤N represents the interference resource allocation of the i-th jammer, where 0≤p ij (t)≤P max , 1≤j≤M,p ij (t) = 0 means that the i-th jammer does not interfere with the j-th link, otherwise it means that the i-th jammer interferes with the j-th link and the interference signal power is p ij (t), and satisfies and sign is the sign function.
[0024] Reward function R: The role of the reward function mechanism in reinforcement learning is to tell the agent the relative goodness of the current behavior, so the reward function can guide the optimization direction of the reinforcement learning method; in the interference resource allocation optimization problem of communication confrontation, the jammer's goal is to make the interference power as small as possible under the premise of achieving the expected symbol error rate, and avoid excessive power to expose the jammer's position. The reward function is defined as:
[0025]
[0026] Among them, ω i is the relative importance coefficient of the ith communication link; sign is the sign function; is the symbol error rate of the ith link; τ0 is the set symbol error rate threshold; P i (t) is the total interference power to the i-th link;
[0027] The goal of the interference resource allocation optimization problem is to maximize the interference effectiveness of the allocation scheme, that is, to maximize the cumulative reward obtained by the interferer over a period of time in the Markov decision process model:
[0028]
[0029] Among them, γ∈[0,1] is the discount factor, which indicates the influence of future returns on the current state; T is the time period of the Markov decision process model;
[0030] The training goal of using reinforcement learning to solve the interference resource allocation optimization problem is to find the optimal strategy π * , achieving return G t Expected maximization. That is:
[0031]
[0032] in, Returns the maximum strategy, E[·] means averaging, G t is the reward function.
[0033] In S3, the objective function of the deep reinforcement learning model is constructed by introducing the maximum policy distribution entropy, as follows:
[0034] As a machine learning method that does not require prior information, deep reinforcement learning adopts a trial-and-error approach to learning. The agent continuously interacts with the environment and takes actions based on the currently learned strategy in the environment. The actions taken will change the state of the environment. The agent then updates and corrects the strategy based on the feedback given by the environment.
[0035] The policy distribution entropy is a parameter that evaluates the randomness of a policy. When the policy distribution entropy is large, it indicates that the policy is more random and has a stronger ability to explore the environment. Sufficient exploration can prevent the policy from falling into the local optimum. Therefore, the concept of policy distribution entropy is introduced into the target reward function to maximize the policy distribution entropy while maximizing the cumulative reward, so that more policies can be explored during the policy optimization process.
[0036] The calculation method of strategy distribution entropy is as follows:
[0037] H t =lg(π φ (a t |s t ))
[0038] H t is the current strategy distribution entropy, is the current strategy;
[0039] By introducing the maximum policy distribution entropy, the objective function of the constructed deep reinforcement model is:
[0040]
[0041] In the formula, argmax means finding a strategy that maximizes the expectation, ρ π is the state-action trajectory distribution formed by strategy π, s t 、a t , r are the state, action and immediate reward at the tth step respectively, E represents the mathematical expectation operation, ∑ t r(s t ,a t ) represents the cumulative reward in a certain period of time, that is, the cumulative interference effectiveness;
[0042] Recursively solve the optimal strategy π * The Q function iteration formula used is:
[0043] Q(s t ,a t )=r t +γE[Q(s t+1 ,a t+1 )-(1-α)H(t)-αH t+1 ]
[0044] The Q function represents the action value in the current state, and α is the update coefficient of the policy network. By giving an initial value, it can be adaptively updated during the learning process.
[0045] In S3, the established interference resource allocation deep reinforcement learning model including the evaluation network, target network and policy network is as follows:
[0046] In the deep reinforcement learning model of interference resource allocation, the evaluation network and the policy network are both neural network models, in which the policy network is a single network structure, which is used to give the current optimal interference resource allocation plan; the evaluation network and the target network use twin networks, that is, two neural networks with the same structure are used to calculate the value of the allocation plan given by the policy network respectively, and compare the values of the allocation plans given by the two policy networks. The larger plan value is used to optimize the parameters of the policy network; after training, the evaluation network converges to the optimal value function, and the policy network converges to the optimal strategy, that is, the best resource allocation plan is obtained.
[0047] In S3, the evaluation network and target network in the interference resource allocation deep reinforcement learning model including the evaluation network, target network and policy network are as follows:
[0048] The twin evaluation network contains two groups of four networks: the first evaluation network Q1, the target network corresponding to the first evaluation network Q1 The second evaluation network Q2 and the target network corresponding to the second evaluation network Q2
[0049] The update rule for the evaluation network is as follows:
[0050] Step 3.1: Calculate the objective function
[0051]
[0052] in, For the target Q network at s t+1 , The corresponding state action value when For the policy network in state s t+1 Reparameterization is a strategy gradient calculation technique that is easy to derive. That is, instead of directly sampling the action using the normal distribution composed of the mean μ and standard deviation σ of the policy network output, random noise ε that satisfies the normal distribution is introduced in the adoption, and the action is generated using the formula a=μ+σ·ε.
[0053] Step 3.2: Define the value loss function
[0054]
[0055] Among them, θ is the evaluation network parameter, is the target network parameter, R t is the current state t Take action a t The instant reward is , and B is the experience replay batch size;
[0056] Step 3.3: Update the evaluation network parameters using gradient descent
[0057]
[0058] in, is the gradient operator, θ1, θ2 are network parameters, α θ is the learning rate;
[0059] Step 3.4: The target network parameters are updated using a single-step soft update method, τ is the flexible update coefficient, and the update method is as follows:
[0060]
[0061] In S3, the strategy network in the interference resource allocation deep reinforcement learning model including the evaluation network, the target network and the strategy network is as follows:
[0062] The policy network uses a single-layer neural network, and the loss function that introduces the maximum policy distribution entropy is as follows:
[0063]
[0064] For the policy network in state s t The reparameterized action outputted below, φ is the policy network parameter, and α is the update coefficient of the policy network;
[0065] The process of updating the policy network parameters using the gradient descent method is as follows:
[0066]
[0067] α φ is the policy network learning rate, is the gradient operator of the policy network.
[0068] In S4, the aforementioned wireless communication confrontation scenario model, Markov decision process model and deep reinforcement learning model are combined and trained to complete the interference strategy allocation. The specific combination and training process is as follows:
[0069] First, the evaluation network, target network, policy network parameters and experience recycling pool are initialized, and then the state space S is initialized. The following actions are repeated during the training process: the Markov decision process model selects action a according to the state and executes it, obtains the immediate reward value of the current action and the next state, and stores the current action, state, and reward value in the experience recycling pool; when the capacity of the experience recycling pool is greater than the specified number, random sampling is performed from the experience recycling pool, the target value is calculated, and the evaluation network parameters and policy network parameters are updated according to the rules until the preset number of training times is reached.
[0070] Compared with the prior art, the present invention has the following advantages:
[0071] 1. A clear adversarial scenario model has been constructed, and the interference resource allocation effect in specific adversarial scenarios is good.
[0072] 2. When constructing the objective function of the deep reinforcement model, the concept of maximum policy distribution entropy is introduced to avoid the decision-making falling into the optimal solution.
[0073] 3. The use of twin evaluation networks can converge to the optimal solution more quickly and make decisions quickly. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Attached Figure 1 It is an interference countermeasure model.
[0075] Attached Figure 2 A framework for allocating methods for reinforcement learning-based disruption policies. DETAILED DESCRIPTION
[0076] The following describes the embodiments of the present application, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0077] A communication countermeasure interference strategy allocation method based on reinforcement learning comprises the following steps: S1: establishing a wireless communication countermeasure scenario model; S2: constructing a Markov decision process model that interacts with the countermeasure scenario; S3: establishing a deep reinforcement learning model for interference resource allocation including an evaluation network, a target network and a strategy network; S4: combining and training the aforementioned wireless communication countermeasure scenario model, the Markov decision process model and the deep reinforcement learning model to complete the interference strategy allocation.
[0078] In S1, the wireless communication confrontation scenario model is established as follows:
[0079] The wireless communication confrontation scenario model is a many-to-many model, that is, the jammer has multiple jammers, and the communicator has multiple communication links for networking communication, as follows:
[0080] The jammer has N jammers, represented by the set n = {1,2,…,N}, and the jammer uses a targeted jamming mode. The communicating party uses the TCP / IP protocol for communication and uses M communication links for networking communication. The set of communication links is represented by m = {1,2,…,M}. These communication links use equal bandwidth channels that do not interfere with each other and are orthogonal, and the relative importance index of each communication link is represented by W = [ω1,ω2,…,ω M ].
[0081] In the wireless communication confrontation scenario model, it is assumed that the jammer has mastered the location of the receivers of each communication link of the enemy through communication reconnaissance and intelligence analysis, and the location of each receiver is considered fixed in the model. For the jammer, the relative importance index of each communication link used by the communication party is unknown. The jammer hopes to reasonably allocate jamming resources under resource-constrained conditions to obtain greater jamming effectiveness.
[0082] In the embodiment, the interference countermeasure scenario constructed is as shown in the attached Figure 1 As shown, the jammer uses a jammer to carry out interference attack on the communication link of the receiver of the jammed party.
[0083] In S2, the constructed Markov decision process model is as follows:
[0084] The reinforcement learning method interacts with the scene and solves the problem by constructing a Markov decision process. The Markov decision process contains elements such as the agent Agent, the state space S, the action space A, the reward function R and the discount factor γ. The Markov decision process is defined as follows:
[0085] Intelligent agent: The jammer formulates a jamming plan through the intelligent jamming engine, and the intelligent jamming engine can guide the reconnaissance aircraft to conduct reconnaissance and guide the jammers to conduct coordinated jamming. Therefore, the intelligent jamming engine can be regarded as an intelligent agent in the Markov decision process.
[0086] State space S: The environment state S(t) represents the allocation scheme of interference resources and the interference effect of the interference scheme at the current moment. S(t) is a (N+1)-row M-column matrix composed of the interference resource allocation matrix X(t) and the interference effect evaluation matrix E(t), that is:
[0087]
[0088] The interference resource allocation matrix is:
[0089] X(t)=[x1(t),x2(t),…,x N (t)] T
[0090] Among them, x i (t) = [c i1 (t) c i2 (t) … c iM (t)], 1≤i≤N represents the jamming target of a single jammer; c ij (t)∈{0,1},c ij (t) = 1 means that the i-th jammer interferes with the j-th communication link, c ij =0 means that the i-th jammer does not interfere with the j-th communication link.
[0091] The interference effect evaluation matrix is:
[0092] E(t)=[τ1(t) τ2(t) … τ M (t)]
[0093] Among them, τ j (t)∈{0,1},τ j (t) = 1 represents the symbol error rate of the jth communication link evaluated by the interferer Reach the preset value τ0, τ j (t)=0 indicates that the symbol error rate does not reach the preset value τ0.
[0094] Action space A: Each jammer can choose to interfere with at most U communication links at time t, and apply a total power of no more than P on the corresponding channels. max The interference signal is , so the interference strategy of the interferer, that is, the interference action, is:
[0095] A(t)=[a1(t) a2(t) … a N (t)] T
[0096] Among them, a i (t) = [p i1 (t) p i2 (t) … p iM (t)], 1≤i≤N represents the interference resource allocation of the i-th jammer, where 0≤p ij (t)≤P max , 1≤j≤M. p ij (t) = 0 means that the i-th jammer does not interfere with the j-th link, otherwise it means that the i-th jammer interferes with the j-th link and the interference signal power is p ij (t), and satisfies and sign is the sign function.
[0097] Reward function R: The role of the reward function mechanism in reinforcement learning is to tell the agent the relative goodness of the current behavior, so the reward function can guide the optimization direction of the reinforcement learning method. In the interference resource allocation optimization problem of communication confrontation, the goal of the jammer is to make the interference power as small as possible while achieving the expected symbol error rate, avoiding excessive power and exposing the jammer's position. The reward function is defined as:
[0098]
[0099] Among them, ω i is the relative importance coefficient of the ith communication link; sign is the sign function; is the symbol error rate of the ith link; τ0 is the set symbol error rate threshold; P i (t) is the total interference power to the i-th link.
[0100] The goal of the interference resource allocation optimization problem is to maximize the interference effectiveness of the allocation scheme, that is, to maximize the cumulative reward obtained by the interferer over a period of time in the Markov decision process model:
[0101]
[0102] Among them, γ∈[0,1] is the discount factor, which indicates the influence of future returns on the current state; T is the time period of the Markov decision process model.
[0103] The training goal of using reinforcement learning to solve the interference resource allocation optimization problem is to find the optimal strategy π * , achieving return G t Expected maximization. That is:
[0104]
[0105] in, Returns the maximum strategy, E[·] means averaging, G t is the reward function.
[0106] In S1, the objective function of the deep reinforcement learning model is constructed by introducing the maximum policy distribution entropy, as follows:
[0107] Deep reinforcement learning is a machine learning method that does not require prior information. It uses trial and error to learn. The intelligent agent continuously interacts with the environment and takes actions based on the current learned strategy in the environment. The actions taken will change the state of the environment. The intelligent agent then updates and corrects the strategy based on the feedback given by the environment.
[0108] The policy distribution entropy is a parameter that evaluates the randomness of a policy. When the policy distribution entropy is large, it indicates that the policy is more random and has a stronger ability to explore the environment. Sufficient exploration can prevent the policy from falling into a local optimum. Therefore, the concept of policy distribution entropy is introduced into the target reward function. The policy distribution entropy is maximized while maximizing the cumulative reward, so that more policies can be explored during the policy optimization process.
[0109] The calculation method of strategy distribution entropy is as follows:
[0110] H t =lg(π φ (a t |s t ))
[0111] H t is the current policy distribution entropy, α is the entropy coefficient, and by giving an initial value, it can be adaptively updated during the learning process, π φ (a t |s t ) is the current strategy.
[0112] By introducing the maximum policy distribution entropy, the objective function of the constructed deep reinforcement model is:
[0113]
[0114] In the formula, argmax means finding a strategy that maximizes the expectation, ρ π is the state-action trajectory distribution formed by strategy π, st 、a t , r are the state, action and immediate reward at the tth step respectively, E represents the mathematical expectation operation, ∑ t r(s t ,a t ) represents the cumulative reward within a certain period of time, that is, the cumulative interference effectiveness.
[0115] Recursively solve the optimal strategy π * The Q function iteration formula used is:
[0116] Q(s t ,a t )=r t +γE[Q(s t+1 ,a t+1 )-(1-α)H(t)-αH t+1 ]
[0117] The Q function represents the action value in the current state, and α is the update coefficient of the policy network. By giving an initial value, it can be adaptively updated during the learning process.
[0118] In S3, the established interference resource allocation deep reinforcement learning model including the evaluation network, target network and policy network is as follows:
[0119] In the interference resource allocation deep reinforcement learning model, both the evaluation network and the policy network are neural network models, where the policy network is a single network structure, used to give the current optimal interference resource allocation plan; the evaluation network and the target network use twin networks, that is, two neural networks with the same structure are used to calculate the value of the allocation plan given by the policy network, and compare the allocation plan values given by the two policy networks. The larger plan value is used to optimize the policy network parameters. After training, the evaluation network converges to the optimal value function, and the policy network converges to the optimal strategy, that is, the optimal resource allocation plan is obtained.
[0120] In S3, the evaluation network and target network in the interference resource allocation deep reinforcement learning model including the evaluation network, target network and policy network are as follows:
[0121] The twin evaluation network contains two groups of four networks: the first evaluation network Q1, the target network corresponding to the first evaluation network Q1 The second evaluation network Q2 and the target network corresponding to the second evaluation network Q2
[0122] The update rule for the evaluation network is as follows:
[0123] Step 3.1: Calculate the objective function
[0124]
[0125] in, For the target Q network at s t+1 , The corresponding state action value. For the policy network in state s t+1 Reparameterization is a strategy gradient calculation technique that is easy to derive. That is, instead of directly sampling the action using the normal distribution composed of the mean μ and standard deviation σ of the policy network output, random noise ε that satisfies the normal distribution is introduced in the sampling, and the action is generated using the formula a=μ+σ·ε.
[0126] Step 3.2: Define the value loss function
[0127]
[0128] Among them, θ is the evaluation network parameter, is the target network parameter, R t is the current state t Take action a t The instant reward is , and B is the experience replay batch size.
[0129] Step 3.3: Update the evaluation network parameters using gradient descent
[0130]
[0131] in, is the gradient operator, θ1, θ2 are network parameters, α θ is the learning rate.
[0132] Step 3.4: The target network parameters are updated using a single-step soft update method, τ is the flexible update coefficient, and the update method is as follows:
[0133]
[0134] In S3, the strategy network in the interference resource allocation deep reinforcement learning model including the evaluation network, the target network and the strategy network is as follows:
[0135] The policy network uses a single-layer neural network, and the loss function that introduces the maximum policy distribution entropy is as follows:
[0136]
[0137] For the policy network in state s t The reparameterized action outputted below, φ is the policy network parameter, and α is the update coefficient of the policy network.
[0138] The process of updating the policy network parameters using the gradient descent method is as follows:
[0139]
[0140] α φ is the policy network learning rate, is the gradient operator of the policy network (a commonly used mathematical operator).
[0141] In S4, the aforementioned wireless communication confrontation scenario model, Markov decision process model and deep reinforcement learning model are combined and trained to complete the interference strategy allocation. The specific combination and training process is as follows:
[0142] First, the evaluation network, target network, policy network parameters and experience recycling pool are initialized. Then the state space S is initialized. The following actions are performed repeatedly during the training process: the Markov decision process model selects action a according to the state and executes it, obtains the immediate reward value and the next state of the current action, and stores the current action, state, and reward value in the experience recycling pool. When the capacity of the experience recycling pool is greater than the specified number, random sampling is performed from the experience recycling pool, the target value is calculated, and the evaluation network parameters and policy network parameters are updated according to the rules. Until the preset number of training times is reached.
[0143] In the embodiment, the interference strategy allocation method framework based on reinforcement learning is constructed as shown in the attached Figure 2 As shown in the figure, by randomly initializing the evaluation network parameters and the policy network parameters, giving the initial decision, and entering the interference environment for policy verification, the obtained policy and its value are put into the experience recycling pool. The policy and its value in the experience recycling pool are sampled and enter the evaluation network and the policy network in turn for parameter update, completing a policy update.
[0144] The training process in the embodiment sets the flexible update coefficient τ = 0.005, the number of interactions per round T = 1000, the experience recovery pool capacity D = 106, the batch sample size B = 256, the discount factor γ = 0.1, and the initial value of the entropy coefficient α = 1. In the embodiment, after training, the obtained suppression strategy can achieve a nearly 95% interference suppression success rate in a single round.
[0145] The invention provides a communication interference strategy allocation method based on reinforcement learning, which is beneficial to enhancing the interference effect, can accelerate the convergence speed of interference strategy allocation, improve the interference level, save time cost, and is beneficial to realizing higher intensity information confrontation, and has important application value for providing interference resource allocation strategy in a short time.
[0146] The above description of the present invention is provided to enable any person of ordinary skill in the art to implement or use the present invention. It will be apparent to those of ordinary skill in the art that various modifications to the present invention are made, and the general principles defined herein may be applied to other variations without departing from the scope of protection of the present invention. Therefore, the present invention is not limited to the examples and designs described herein, but is consistent with the broadest range of principles and novel features disclosed herein.
Claims
1. A communication anti-interference strategy allocation method based on reinforcement learning, characterized in that: The following steps are involved: S1: Establish a wireless communication adversarial scenario model; S2: Construct a Markov decision process model that interacts with the adversarial scenario; S3: Establish a deep reinforcement learning model for interference resource allocation that includes an evaluation network, a target network, and a strategy network; S4: Combine and train the aforementioned wireless communication adversarial scenario model, the Markov decision process model, and the deep reinforcement learning model to complete the interference strategy allocation.
2. According to the method for allocating communication anti-interference strategies based on reinforcement learning in claim 1, it is characterized in that: In S1, the wireless communication confrontation scenario model is established as follows: The wireless communication confrontation scenario model is a many-to-many model, that is, the jammer has multiple jammers, and the communicator has multiple communication links for networking communication, as follows: The jammer has N jammers, represented by the set n = {1, 2, ..., N}, and the jammer adopts the aiming jamming mode; the communicator uses the TCP / IP protocol for communication and uses M communication links for networking communication. The set of communication links is represented by m = {1, 2, ..., M}. These communication links use equal bandwidth channels that do not interfere with each other and are orthogonal, and the relative importance index of each communication link is represented by W = [ω1, ω2, ..., ω M ]; In the wireless communication confrontation scenario model, it is assumed that the jammer has mastered the locations of the receivers of each communication link of the enemy through communication reconnaissance and intelligence analysis, and the model assumes that the locations of each receiver are fixed; for the jammer, the relative importance index of each communication link used by the communicator is unknown; the jammer hopes to reasonably allocate jamming resources under resource-constrained conditions to obtain greater jamming effectiveness.
3. According to the method for allocating communication anti-interference strategies based on reinforcement learning in claim 1, it is characterized in that: In S2, the constructed Markov decision process model is as follows: The reinforcement learning method interacts with the scene and solves the problem by constructing a Markov decision process; the Markov decision process includes the agent Agent, the state space S, the action space A, the reward function R and the discount factor γ elements; the Markov decision process is defined as follows: Agent: The jammer formulates a jamming plan through the intelligent jamming engine, and the intelligent jamming engine guides the reconnaissance aircraft to conduct reconnaissance and guides the jammers to conduct coordinated jamming. Therefore, the intelligent jamming engine is regarded as an agent in the Markov decision process; State space S: The environment state S(t) represents the allocation scheme of interference resources and the interference effect of the interference scheme at the current moment. S(t) is a (N+1)-row M-column matrix composed of the interference resource allocation matrix X(t) and the interference effect evaluation matrix E(t), that is: The interference resource allocation matrix is: X(t)=[x1(t),x2(t),…,x N (t)] T Among them, x i (t) = [c i1 (t) c i2 (t)…c iM (t)], 1≤i≤N represents the jamming target of a single jammer; c ij (t)∈{0,1},c ij (t) = 1 means that the i-th jammer interferes with the j-th communication link, c ij =0 means that the i-th jammer does not interfere with the j-th communication link; The interference effect evaluation matrix is: E(t)=[τ1(t) τ2(t)…τ M (t)] Among them, τ j (t)∈{0,1},τ j (t) = 1 represents the symbol error rate of the jth communication link evaluated by the interferer Reach the preset value τ0, τ j (t) = 0 means that the symbol error rate does not reach the preset value τ0; Action space A: Each jammer at time t can jam at most U communication links and exert a total power of no more than P on the corresponding channels. max The interference signal is , so the interference strategy of the interferer, that is, the interference action, is: A(t)=[a1(t) a2(t)…a N (t)] T Among them, a i (t) = [p i1 (t) p i2 (t)…p iM (t)], 1≤i≤N represents the interference resource allocation of the i-th jammer, where 0≤p ij (t)≤P max , 1≤j≤M,p ij (t) = 0 means that the i-th jammer does not interfere with the j-th link, otherwise it means that the i-th jammer interferes with the j-th link and the interference signal power is p ij (t), and satisfies and sign is the sign function. Reward function R: The role of the reward function mechanism in reinforcement learning is to tell the agent the relative goodness of the current behavior, so the reward function can guide the optimization direction of the reinforcement learning method; in the interference resource allocation optimization problem of communication confrontation, the jammer's goal is to make the interference power as small as possible under the premise of achieving the expected symbol error rate, and avoid excessive power to expose the jammer's position. The reward function is defined as: Among them, ω i is the relative importance coefficient of the ith communication link; sign is the sign function; is the symbol error rate of the ith link; τ0 is the set symbol error rate threshold; P i (t) is the total interference power to the i-th link; The goal of the interference resource allocation optimization problem is to maximize the interference effectiveness of the allocation scheme, that is, to maximize the cumulative reward obtained by the interferer over a period of time in the Markov decision process model: Among them, γ∈[0,1] is the discount factor, which indicates the influence of future returns on the current state; T is the time period of the Markov decision process model; The training goal of using reinforcement learning to solve the interference resource allocation optimization problem is to find the optimal strategy π * , achieving return G t Expected maximization. That is: in, Returns the maximum strategy, E[·] means averaging, G t is the reward function.
4. According to claim 3, a communication anti-interference strategy allocation method based on reinforcement learning is characterized in that: In S3, the objective function of the deep reinforcement learning model is constructed by introducing the maximum policy distribution entropy, as follows: As a machine learning method that does not require prior information, deep reinforcement learning adopts a trial-and-error approach to learning. The agent continuously interacts with the environment and takes actions based on the currently learned strategy in the environment. The actions taken will change the state of the environment. The agent then updates and corrects the strategy based on the feedback given by the environment. The policy distribution entropy is a parameter that evaluates the randomness of a policy. When the policy distribution entropy is large, it indicates that the policy is more random and has a stronger ability to explore the environment. Sufficient exploration can prevent the policy from falling into the local optimum. Therefore, the concept of policy distribution entropy is introduced into the target reward function to maximize the policy distribution entropy while maximizing the cumulative reward, so that more policies can be explored during the policy optimization process. The calculation method of strategy distribution entropy is as follows: H t =lg(π φ (from t |s t )) H t is the current policy distribution entropy, π φ (a t |s t ) is the current strategy; By introducing the maximum policy distribution entropy, the objective function of the constructed deep reinforcement model is: In the formula, argmax means finding a strategy that maximizes the expectation, ρ π is the state-action trajectory distribution formed by strategy π, s t 、a t , r are the state, action and immediate reward at the tth step respectively, E represents the mathematical expectation operation, ∑ t r(s t , a t ) represents the cumulative reward in a certain period of time, that is, the cumulative interference effectiveness; Recursively solve the optimal strategy π * The Q function iteration formula used is: Q(s t ,a t )=r t +γE[Q(s t+1 ,a t+1 )-(1-α)H(t)-αH t+1 ] The Q function represents the action value in the current state, and α is the update coefficient of the policy network. By giving an initial value, it can be adaptively updated during the learning process.
5. According to a method for allocating communication anti-interference strategies based on reinforcement learning in claim 1, it is characterized in that: In S3, the established interference resource allocation deep reinforcement learning model including the evaluation network, target network and policy network is as follows: In the deep reinforcement learning model for interference resource allocation, both the evaluation network and the policy network are neural network models, where the policy network is a single network structure, which is used to give the current optimal interference resource allocation solution; The evaluation network and the target network use twin networks, that is, two neural networks with the same structure, to calculate the value of the allocation plan given by the policy network respectively, and compare the values of the allocation plans given by the two policy networks. The larger plan value is used to optimize the parameters of the policy network. After training, the evaluation network converges to the optimal value function, and the policy network converges to the optimal strategy, that is, the best resource allocation plan is obtained.
6. The communication anti-interference strategy allocation method based on reinforcement learning according to claim 1 is characterized in that: In S3, the evaluation network and target network in the interference resource allocation deep reinforcement learning model including the evaluation network, target network and policy network are as follows: The twin evaluation network contains two groups of four networks: the first evaluation network Q1, the target network corresponding to the first evaluation network Q1 The second evaluation network Q2 and the target network corresponding to the second evaluation network Q2 The update rule for the evaluation network is as follows: Step 3.1: Calculate the objective function in, For the target Q network at s t+1 , The corresponding state action value when For the policy network in state s t+1 Reparameterization is a strategy gradient calculation technique that is easy to derive. That is, instead of directly sampling the action using the normal distribution composed of the mean μ and standard deviation σ of the policy network output, random noise ε that satisfies the normal distribution is introduced in the adoption, and the action is generated using the formula a=μ+σ·ε. Step 3.2: Define the value loss function Among them, θ is the evaluation network parameter, is the target network parameter, R t is the current state t Take action a t The instant reward is , and B is the experience replay batch size; Step 3.3: Update the evaluation network parameters using gradient descent in, is the gradient operator, θ1, θ2 are network parameters, α θ is the learning rate; Step 3.4: The target network parameters are updated using a single-step soft update method, τ is the flexible update coefficient, and the update method is as follows:
7. The communication anti-interference strategy allocation method based on reinforcement learning according to claim 1 is characterized in that: In S3, the strategy network in the interference resource allocation deep reinforcement learning model including the evaluation network, the target network and the strategy network is as follows: The policy network uses a single-layer neural network, and the loss function that introduces the maximum policy distribution entropy is as follows: For the policy network in state s t The reparameterized action outputted below, φ is the policy network parameter, and α is the update coefficient of the policy network; The process of updating the policy network parameters using the gradient descent method is as follows: α φ is the policy network learning rate, is the gradient operator of the policy network.
8. The communication anti-interference strategy allocation method based on reinforcement learning according to claim 1 is characterized in that: In S4, the aforementioned wireless communication confrontation scenario model, Markov decision process model and deep reinforcement learning model are combined and trained to complete the interference strategy allocation. The specific combination and training process is as follows: First, the evaluation network, target network, policy network parameters and experience recycling pool are initialized, and then the state space S is initialized. The following actions are repeated during the training process: the Markov decision process model selects action a according to the state and executes it, obtains the immediate reward value of the current action and the next state, and stores the current action, state, and reward value in the experience recycling pool; when the capacity of the experience recycling pool is greater than the specified number, random sampling is performed from the experience recycling pool, the target value is calculated, and the evaluation network parameters and policy network parameters are updated according to the rules until the preset number of training times is reached.
Citation Information
Patent Citations
Deep reinforcement learning communication interference resource allocation method fused with noise network
CN115866760A
Dynamic interference power distribution method for incomplete perception
CN118678451A
Unmanned aerial vehicle cluster interference resource allocation method based on pre-training attention encoder
CN118764958A
Method and system for developing battlefield ai taxonomy
US20230169370A1