Deep reinforcement learning aided target threat assessment model and jamming beam allocation method

By constructing a target threat assessment model and an interference beam allocation method, and using the TSM function and DSAC algorithm to reconstruct the target threat matrix, the high-dimensional difficulty in multi-target interference resource allocation is solved, efficient interference beam resource allocation is achieved, and the interference effectiveness of the electronic countermeasure system is improved.

CN119342585BActive Publication Date: 2025-10-24CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411449576.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2025-10-24
Estimated Expiration
2044-10-17

AI Technical Summary

Technical Problem

The existing technology has problems of slow convergence and poor scalability in multi-target interference resource allocation, and lacks a target threat assessment link, making it difficult to meet the needs of cognitive electronic countermeasures systems for intelligent allocation of large-scale interference resources.

Method used

A collaborative jamming model for multiple jamming devices to counter enemy UAV formations is constructed. The target threat assessment link is introduced. The target threat assessment matrix based on the TSM function and the DSAC algorithm are used to reconstruct the target threat assessment matrix, generate the jamming beam resource allocation matrix, and realize the reasonable allocation of high-dimensional jamming beam resources with the assistance of deep reinforcement learning.

Benefits of technology

It effectively solves the problem of difficult allocation of high-dimensional interference beam resources, realizes efficient interference resource allocation in a dynamically changing environment, and improves the interference effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119342585B_ABST
    Figure CN119342585B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of deep reinforcement learning assisted target threat degree evaluation model and interference beam allocation method, belong to wireless mobile communication technical field.The present application includes the following steps:S1: in wartime communication confrontation scene, construct the cooperative interference model of multiple interference devices to enemy unmanned aerial vehicle formation;S2: analyze the constraint condition of interference beam allocation process;S3: introduce target threat degree evaluation link, construct interference beam resource allocation optimization problem;S4: based on tactical importance plot (TSM) function constructs target threat degree evaluation model, generates target threat degree evaluation matrix;S5: TSM function is improved using discrete flexible actor-critic (DSAC) algorithm, reconstructs target threat degree evaluation matrix;S6: according to target threat degree evaluation matrix, executes interference beam allocation.The present application can realize the reasonable allocation of interference beam in " multiple to multiple " wartime communication confrontation scene, to effectively improve the utilization of interference resource.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of wireless mobile communication, and relates to a target threat degree evaluation model assisted by deep reinforcement learning and a method for allocating interference beams BACKGROUND

[0002] The deep integration of artificial intelligence (AI) technology, especially machine learning (ML), and electronic warfare (EW) promotes the leapfrog development of EW from "manual cognition" to "machine cognition". In the war environment where intelligent electronic equipment threats are increasingly severe, the struggle for control of the electromagnetic spectrum between the two sides of cognitive electronic warfare is becoming more intense, and resource constraints have become the main constraint on the effectiveness of the confrontation between the two sides. Therefore, whether limited interference resources can be efficiently and reasonably allocated has become the key to maximizing the effectiveness of electronic warfare jamming.

[0003] Deep reinforcement learning (DRL) combines deep learning (DL) and reinforcement learning (RL), giving the agent strong autonomous decision-making ability and abstract representation ability of the environment, and has the advantages of fast solving speed, support for multi-dimensional decision-making, and strong generalization ability. With the continuous progress of technology, the electronic countermeasure environment is becoming increasingly complex, and "multi-to-multi" cluster combat has gradually become the mainstream of electronic warfare in wartime. Interference equipment is required to have the ability to interfere with multiple targets simultaneously.

[0004] Some scholars have studied the resource allocation problem of multi-target jamming in multi-beam ground-to-air radar jamming system by comprehensively considering the individual and overall jamming effects (Reference: Cui Z M, Peng S L, Ren M Q, et al. Multi-target jamming decision research of multi-beam jamming system[J]. Firepower and Command Control, 2021, 46(12): 149-155.); some research works have discussed the joint scheduling problem of jamming beams and power (Reference: ZHANG D, SUN J, YI W, et al. Joint jamming beam and power scheduling for suppressing netted radar system[C] / / 2021 IEEE Radar Conference (RadarConf21). IEEE, 2021: 1-6.). However, these researches mainly rely on conventional multi-objective optimization theory in the design of multi-target jamming resource allocation scheme; but due to the complexity of high-dimensional jamming resource allocation, these schemes often face challenges such as slow convergence speed and poor scalability, making it difficult to meet the demand of cognitive electronic countermeasure system for intelligent allocation of large-scale jamming resources, and related work simplifies or lacks the target threat degree evaluation link when solving the jamming resource allocation problem. Target threat degree evaluation, as a core component of command, control, decision-making and other links in complex electronic countermeasure process in wartime, is a prerequisite and important basis for realizing jamming resource allocation. SUMMARY

[0005] The present application aims to provide a target threat degree evaluation model and jamming beam allocation method assisted by deep reinforcement learning, which can realize the reasonable allocation of high-dimensional jamming beam resources by introducing the target threat degree evaluation link. The technical solution of the present application is as follows:

[0006] S1: In the communication countermeasure scene in wartime, a cooperative jamming model of multiple jamming devices against enemy unmanned aerial vehicle formation is constructed;

[0007] S2: Analyze the resource constraint conditions of the beam resource allocation process;

[0008] S3: Introduce the target threat degree evaluation link, construct the jamming beam resource allocation optimization problem, and guide the reasonable allocation of jamming beam resources;

[0009] S4: Based on the Tactical Significance Map (TSM) function, a target threat degree evaluation model is constructed to generate a target threat degree evaluation matrix;

[0010] S5: The TSM function is improved by using the Discrete Soft Actor-Critic (DSAC) algorithm, and the target threat degree evaluation matrix is reconstructed.

[0011] S6: Based on the target threat degree evaluation matrix, the interference beam resource allocation strategy is executed to generate an interference beam resource allocation matrix.

[0012] Further, in S1, a "many-to-many" wartime communication confrontation scenario is considered, in which M jamming devices of the interference party implement suppressive jamming on the unmanned aerial vehicle formation composed of N unmanned aerial vehicles commanded by the enemy early warning aircraft; N e ={1, 2,..., N} and M e ={1, 2,..., M} represent sets composed of N unmanned aerial vehicles in the unmanned aerial vehicle formation and M jamming devices, respectively. To avoid jamming, the unmanned aerial vehicle formation is in dynamic change in speed and direction during flight to the interference party position; the jamming device can generate multiple beams to simultaneously jam multiple unmanned aerial vehicle targets.

[0013] Further, in S2, the beam pointing vector of the jamming device m is b m =[b m,1 ,b m,2 ,...,b m,N ] T , where b m,n ∈{0, 1} is a beam pointing variable, [·] T represents the transpose operation, and m, n represent the index numbers of the jamming device and the unmanned aerial vehicle target. When the jamming device m ∈ M e allocates a beam to the unmanned aerial vehicle target n ∈ N e to implement jamming, b m,n =1, otherwise b m,n =0; the beam resource allocation matrix B=[b1, b2,..., b M ] T is defined. Considering that the unmanned aerial vehicle targets in the unmanned aerial vehicle formation are relatively dispersed in space, a single beam cannot simultaneously cover multiple targets, it is assumed that a beam of each jamming device can only jam a single unmanned aerial vehicle target; at the same time, due to the performance constraints of the jamming device, a single jamming device can simultaneously jam at most I unmanned aerial vehicle targets:

[0014]

[0015] To ensure that the interference beam resources are reasonably utilized, it is limited that each unmanned aerial vehicle is jammed by at most U beams:

[0016]

[0017] Further, in S3, when allocating jamming resources, the difference in threat degree of different unmanned aerial vehicle targets needs to be considered. According to the evaluation result of the jamming device m, the target threat degree evaluation vector w m =[w m,1 ,wm,2,... ,w m,N ] T wherein 0≤w m,n ≤1; the target threat assessment matrix W is defined as W=[w1,w2,...,wn]T based on the assessment results of all the interference devices. M ] T The average threat of the UAV target n is defined as The interference beam resource allocation problem of multiple interference devices against the UAV formation can be expressed as:

[0018]

[0019] In the above formula, P1 is the interference beam allocation problem, max represents maximization, and P1 represents maximizing the sum of the average threats of all the UAV targets by solving the beam resource allocation matrix B under the constraint conditions C1 and C2.

[0020] Further, in the S4, the target threat assessment is performed based on the TSM function, which can be expressed as:

[0021]

[0022] wherein, is the state vector of the UAV target n at the current time, [x n ,y n ] T and are the position vector and the velocity vector of the UAV target n, respectively, the reference point coordinates are h c =[x n ,y n ] T , exp is the exponential function, represents the distance between the UAV target n and the reference point, k0 and m0 are preset constants, is the absolute velocity of the UAV target n, θ c,n is the heading angle of the UAV target n relative to the reference point, and its expression is arccos represents the inverse cosine function; all the interference devices are taken as the reference points, and the target threat assessment is performed against the UAV targets to generate the target threat assessment matrix W.

[0023] Further, in S5, when the target distance of the UAV is far from the jammer, the target threat assessment result based on the TSM function will degenerate into the assessment result based on a single jamming device. To solve this problem, the target threat assessment matrix reconstruction problem is modeled as a Markov Decision Process (MDP), including state space, action space and reward function. The motion state information of the UAV target n at time t for M jamming devices is s n,b,t = [o 1,b,t , o 2,b,t , ..., o M,b,t ], wherein subscript b represents beam, t represents time, o m,b.t = [R m,n,t , v n,t , θ m,n,t ], containing the distance R m,n,t , velocity v n,t and heading angle θ m,n,t of the UAV target n to the jamming device m at time t. Based on s n,b,t , the state of MDP is defined as s b,t = [s 1,b,t , s 2,b,t , ..., s N,b,t ] ∈ S b , wherein S b is the state space; the action space is M e ; the reward function is based on the TSM function, when the jamming device is selected to interfere with the UAV target n, if the threat assessment value w m,n of the selected jamming device m is higher than all assessment results w n = [w m,1 , w m,2 , ..., w m,N ] T , the higher the reward value obtained is.

[0024] Further, in S5, the target threat assessment matrix is reconstructed by using the DSAC algorithm. The reconstruction process of the target threat assessment matrix includes the following steps:

[0025] S71: initialize the number of jamming devices M, the number of UAV targets N, the total number of training episodes max_episode, the total interaction time step max_t of each episode, the parameters of the policy (Actor) network, the evaluation (Critic) network and the target evaluation (Critic) network, the experience replay pool, the threshold Φ b , let B = 0, W = 0;

[0026] S72: If the number of training episodes exceeds max_episode, end the training, execute S76, otherwise initialize max_t, execute S73;

[0027] S73: If the number of interaction time steps exceeds max_t, execute S72, otherwise execute S74;

[0028] S74: For the UAV target n, the agent Actor network inputs the state s n,b,t , outputs the action a n,b,t , and obtains the reward r n,b,t , the state is transferred to s n,b,t+1 , the experience sample (s n,b,t , a n,b,t , r n,b,t , s n,b,t+1 ) is stored in the experience replay pool;

[0029] S75: If the number of experience replay pool samples is greater than the threshold Φ b , randomly sample a batch of samples, train the Actor and Critic networks, otherwise execute S73;

[0030] S76: Save the DSAC model;

[0031] Further, in the S6, based on the target threat degree evaluation matrix, the jamming beam resource allocation strategy is executed to generate a jamming beam resource allocation matrix, including the following steps:

[0032] S81: For each UAV target n, the agent Actor network inputs s n,b,t , outputs the probability vector ξ n , and saves it, and reconstructs the target threat degree evaluation matrix W = [ξ1, ξ2,..., ξ N ];

[0033] S82: Loop through W in descending order, select the position with the maximum threat degree evaluation value in W and set it to 0, map the value of the corresponding position in the jamming beam allocation matrix B to 1, and the mapping process needs to satisfy the C1 and C2 constraint conditions at the same time, until each UAV target is irradiated by a jamming beam;

[0034] S83: Based on S82, repeat the loop mapping process, when the number of UAVs allocated by each jamming device reaches the maximum number of UAV targets I that can be simultaneously jammed, the beam resource allocation process ends;

[0035] S84: Output the jamming beam resource allocation matrix B

[0036] The beneficial effects of the present application are that the present application proposes a deep reinforcement learning assisted communication countermeasure interference resource allocation method for the problem of high-dimensional interference beam resource allocation difficulty caused by environmental dynamic change characteristics and the increase of the number of communication countermeasure members in the "many-to-many" communication countermeasure scene. First, a cooperative jamming model of multiple jamming devices against enemy unmanned aerial vehicle formation is established; then, considering the constraints of limited interference resources and interference device performance, a target threat degree evaluation link is introduced, and an interference beam resource allocation optimization problem is established; then, the target threat degree evaluation matrix is constructed based on the TSM function, and the target threat degree evaluation matrix is reconstructed using DSAC for the shortcomings of the TSM function; finally, the DSAC network model is saved, the interference beam allocation matrix is decided by the Actor network, and the reasonable allocation of high-dimensional interference beam resources is effectively realized. BRIEF DESCRIPTION OF DRAWINGS

[0037] In order to make the purpose, technical scheme and advantages of the present application clearer, the preferred detailed description of the present application will be described below in combination with the drawings, in which:

[0038] Figure 1 Cooperative jamming model of multiple jamming devices against enemy unmanned aerial vehicle formation Figure 2 DSAC algorithm network structure diagram

[0039] Figure 3 Overall flow chart of interference beam allocation DETAILED DESCRIPTION

[0040] The embodiment of the present application proposes a deep reinforcement learning assisted target threat degree evaluation model and interference beam allocation method, which can effectively realize the reasonable allocation of high-dimensional interference beam resources.

[0041] The specific steps of the method are as follows:

[0042] S1: In the wartime communication countermeasure scene, a cooperative jamming model of multiple jamming devices against enemy unmanned aerial vehicle formation is constructed. In the model, the jamming side includes M jamming devices, which implement suppressive jamming on the unmanned aerial vehicle formation composed of N unmanned aerial vehicles commanded by the enemy early warning aircraft; N e ={1,2,…,N} and M e ={1,2,…,M} represent the sets composed of N unmanned aerial vehicles in the unmanned aerial vehicle formation and M jamming devices respectively; in order to avoid jamming, the unmanned aerial vehicle formation dynamically changes the speed and direction during flight to the jamming side; the jamming device can generate multiple beams to simultaneously jam multiple unmanned aerial vehicle targets.

[0043] S2: The beam pointing vector of the jamming device m is b m =[b m,1 ,b m,2 ,…,b m,N ]T where b m,n ∈{0,1} is the beam pointing variable, [·] T denotes the transpose operation, and the subscripts m,n denote the index number of the jammer and the UAV target. When the jammer m∈M e allocates a beam to the UAV target n∈N e b m,n = 1 when implementing jamming, otherwise b m,n = 0; define the beam resource allocation matrix B = [b1, b2, …, b M ] T Considering that the UAVs in the formation are distributed in space, a single beam cannot cover multiple targets, it is assumed that a beam of each jammer can only interfere with a single UAV target; at the same time, due to the performance constraints of the jammer, a single jammer can simultaneously interfere with at most I UAV targets:

[0044]

[0045] To ensure that the jamming beam resources are reasonably utilized, each UAV is limited to being interfered by at most U beams:

[0046]

[0047] S3: In communication confrontation, the jamming effect is closely related to the jamming resource allocation method. In actual communication confrontation, the greater the threat degree of the target, the more attention it should receive, therefore, when allocating jamming resources, the difference in threat degree of different UAV targets needs to be considered. Therefore, according to the evaluation results of the jammer m, the target threat degree evaluation vector w m = [w m,1 , w m,2 , …, w m,N ] T is defined, where 0≤w m,n ≤1; further, the target threat degree evaluation matrix W = [w1, w2, …, w M ] T is defined by integrating the evaluation results of all jammers.

[0048] In wartime scenarios, to ensure the survivability of each jammer, each jammer is usually deployed in a distributed manner, so there is a certain difference in the threat degree evaluation of different jammers to different targets. Therefore, to fully exploit the geographical advantage of the jammer, for any UAV target, the jammer that evaluates the target threat degree as high as possible should implement jamming on the target, so the jamming beam allocation can be determined according to the target threat degree evaluation matrix. Therefore, the average threat degree of the UAV target n is defined as The jamming beam allocation problem of multiple jammers to the UAV formation can be represented as:

[0049]

[0050] In the above formula, P1 is an interference beam allocation problem, max represents maximization, and P1 represents maximizing the sum of the average threat degrees of all unmanned aerial vehicle targets by solving the beam resource allocation matrix B under the constraint conditions C1 and C2.

[0051] S4: A dynamic assessment model of target threat degree is established based on the TSM in consideration of the motion state difference of different unmanned aerial vehicle targets. The TSM function can be expressed as:

[0052]

[0053] wherein, is the state vector of the unmanned aerial vehicle target n at the current moment, [x n ,y n ] T and represent the coordinates and the speed of the unmanned aerial vehicle target n, respectively, the reference point coordinates are h c = [x c ,y c ] T , exp is an exponential function, represents the distance between the unmanned aerial vehicle target n and the reference point, k0 and m0 represent preset constants, is the absolute speed of the unmanned aerial vehicle target n, and θ c,n is the heading angle of the unmanned aerial vehicle target n relative to the reference point, and its expression is arccos represents an inverse cosine function; all interference devices are taken as reference points, and the target threat degree assessment is performed on the unmanned aerial vehicle targets to generate a target threat degree assessment matrix W

[0054] S5: The target threat degree dynamic assessment model based on the TSM function assesses the target threat degree based on the motion state of the unmanned aerial vehicle target with different interference devices as reference. When the unmanned aerial vehicle target is far away from the interference device, the motion state information of the same unmanned aerial vehicle target to different interference devices is similar. Therefore, in the target threat degree assessment matrix W determined by the TSM function, the threat degree values of the same unmanned aerial vehicle target are very small between different interference devices. At this time, the assessment results in W will degenerate into the threat degree assessment based on a single interference device.

[0055] To solve the above problems, the DSAC is used to reconstruct the target threat assessment matrix, the motion state of all interference devices to the same UAV target is taken as the input of the Actor network, and the target threat assessment result based on the TSM is used to design the reward function. Finally, the output probability vector of the output layer of the Actor network (which is a fully connected layer and outputs a probability vector with a sum of 1) is taken as the new threat assessment vector, that is, each probability value is taken as the new threat assessment value of each interference device to the UAV target. In this way, the threat assessment difference between different interference devices can be amplified, and a new target threat assessment matrix W is constructed. The target threat assessment problem is modeled as a Markov decision process (MDP), including state space, action space and reward function.

[0056] The motion state information of the UAV target n to the M interference devices at time t is s n,b,t = [o 1,b,t ,o 2,b,t ,...,o M,b,t ], where b represents the beam, t represents the time, o m,b.t = [R m,n,t ,v n,t ,θ m,n,t ] contains the distance R m,n,t , the speed v n,t and the heading angle θ m,n,t of the UAV target n to the interference device m at time t. Based on s n,b,t , the state of the MDP is defined as s b,t = [s 1,b,t ,s 2,b,t ,...,s N,b,t ] ∈ S b , where S b is the state space.

[0057] The action space is M e , and the action a n,b,t ∈ M e represents selecting one of the M interference devices to perform interference on the UAV target n; from the target threat assessment matrix W determined based on the TSM, the nth column is selected for the interference target n, that is, w n = [w 1,n ,w 2,n ,...,w M,n ] T , and the reward function is designed as

[0058]

[0059] where w n [a n,b,t ] represents the interference device selected by a nGet action a n,b,t The target threat assessment value w of the corresponding jammer m against the drone target n m,n ,sorted(w n ) represents the vector w n Sort the elements in from small to large, idx(w n [a n,b,t ],sorted(w n )) is used to obtain w m,n After sorting the vector w n The position index in (assuming the position order is (1, 2, ..., M)). This formula shows that when a jammer is selected to jam the drone target n, if the threat assessment value w of the selected jammer is m,n In all evaluation results w n The higher the value, the higher the reward value; define the joint reward r b,t =[r 1,b,t ,r 2,b,t ,...,r N,b,t ], which contains the rewards for N drone targets.

[0060] The DSAC algorithm not only maximizes expected returns and policy entropy simultaneously, but also reduces sensitivity to model and estimation errors, and has the advantages of fast, comprehensive, and stable training. The optimal joint policy solution is constructed as shown in the following formula:

[0061]

[0062] Among them, π * is the optimal joint strategy, argmax represents the independent variable value corresponding to the maximum value of the function, Indicates the expected value, ρ π Represents a state-action pair (s b,t ,a b,t ) corresponds to the distribution function, γ is the discount factor, That is, the strategy π in state s b,t The entropy of , log represents the logarithm, and the entropy temperature coefficient α determines the relative importance of the reward value and the entropy term, thereby controlling the randomness of the optimal strategy.

[0063] To improve the stability of training and reduce the estimation bias of target Q value, this paper adopts the same structure of Critic network as the evaluation (Critic) network, and updates the parameters of Critic network to the target Critic network regularly; in addition, to further alleviate the overestimation problem of Q value, this paper adopts a double neural network structure in Critic network and target Critic network, that is, each network contains two neural networks with the same structure, and in the training process, the Critic network and the target Critic network each select the smaller one of the two neural networks to update the network parameters, so as to optimize the performance and ensure the accuracy of estimation.

[0064] To evaluate the state s n,b,t of the target UAV n at time t, define the soft value function

[0065]

[0066] Where θ i is the Critic network parameter, is the action value function corresponding to a n,b,t , -log(π(a n,b, t|s n,b,t )) represents the policy entropy, and φ is the Actor network parameter.

[0067] After performing action a n,b,t in state s n,b,t , the update process of the target Q value can be represented as

[0068]

[0069] Where y(a n,b,t ,s n,b,t+1 ) is the target Q value, is the action value corresponding to action a n,b,t+1 in state s n,b,t+1 of the target Critic network, is the parameter of the target Critic network, and min represents the minimum value.

[0070] To effectively evaluate the performance of the executed action, define the loss function of the Critic network

[0071]

[0072] Where, is the experience replay pool.

[0073] In the policy evaluation stage, the Critic network parameters are updated by the stochastic gradient descent method

[0074]

[0075] where λ Q is the learning rate of the Critic network's loss function, denotes the derivation of the parameter θ Q in the J i (θ i ) function;

[0076] In the policy improvement stage, the policy update is achieved by minimizing the Kullback-Leibler (KL) divergence

[0077]

[0078] where Π represents the set of candidate policies, π(·|s n,b,t ) represents the policy to be optimized, denotes the Q function under the current policy, is the partition function of the current policy, and after simplification, the loss function of the Actor network is as follows

[0079]

[0080] The Actor network parameters are updated using the stochastic gradient descent method

[0081]

[0082] where λ φ is the learning rate of the Actor network's loss function, denotes the derivation of the parameter φ in J π (φ);

[0083] By setting the target entropy, the gradient descent method is used to adaptively update α, and the loss function of the entropy temperature coefficient is defined

[0084]

[0085] where, is the target entropy, which is related to the dimension of the agent's action;

[0086] The entropy temperature coefficient α is updated by training J(α)

[0087]

[0088] where λ α is the learning rate of the entropy temperature coefficient's loss function, denotes the derivation of the parameter α in J(α);

[0089] Finally, the parameters of the target Critic network are updated in a flexible manner

[0090]

[0091] wherein τ is a flexible update coefficient.

[0092] The reconstruction process of the target threat assessment matrix includes the following steps:

[0093] S71: initialize the number of interference devices M, the number of UAV targets N, the total number of training episodes max_episode, the total interaction time step max_t of each episode, the parameters of the Actor, Critic network and target Critic, the experience replay pool, and the threshold Φ b , let B=0, W=0;

[0094] S72: if the number of training episodes exceeds max_episode, end the training and execute S76, otherwise initialize max_t and execute S73;

[0095] S73: if the number of interaction time steps exceeds max_t, execute S72, otherwise execute S74;

[0096] S74: for the UAV target n, the agent Actor network inputs s n,b,t , outputs the action a n,b,t and obtains the reward r rn,b,t , the state is transferred to s n,b,t+1 , and the experience sample (s n,b,t , a n,b,t , r n,b,t , s n,b,t+1 ) is stored in the experience replay pool;

[0097] S75: if the number of experience replay pool samples is greater than the threshold Φ b , randomly extract a batch of experience samples, train the Actor and Critic networks, otherwise continue to execute S73;

[0098] S76: save the DSAC model

[0099] S6: based on the target threat assessment matrix, execute the jamming beam resource allocation strategy to generate a jamming beam resource allocation matrix, specifically including the following steps:

[0100] S81: for each UAV target, the agent Actor network inputs n and outputs the probability vector ξ n and saves it, and reconstructs the target threat assessment matrix W=[ξ1,ξ2,...,ξ N ]

[0101] S82: Circulatingly traverse W, select the position with the maximum threat degree evaluation value in W in descending order, set it to 0, and map the value of the corresponding position in the interference beam allocation matrix B to 1, which needs to meet the C1 and C2 constraint conditions at the same time in the mapping process, until each UAV target is irradiated by an interference beam;

[0102] S83: On the basis of S82, repeat the circulation and mapping process, when the number of UAVs allocated to each interference device reaches the maximum number of UAV targets I that can be interfered simultaneously, the interference beam allocation process ends.

[0103] S84: Output the interference beam allocation matrix B.

Claims

1. A target threat assessment model assisted by deep reinforcement learning and a jamming beam resource allocation method, characterized in that: The method comprises the following steps: S1: in the communication confrontation scene in wartime, a cooperative jamming model of multiple jamming devices against enemy unmanned aerial vehicle formation is constructed; S2: Analyzing the resource constraint condition of the beam resource allocation process: the beam pointing vector of the interference device m is b m = [b m,1 , b m,2 ,..., b m,N ] T , where b m,n ∈ {0, 1} is the beam pointing variable, [·] T represents the transpose operation, m, n represent the index numbers of the interference device and the UAV target; the interference device m ∈ M e allocates the beam to the UAV target n ∈ N e When implementing interference, b m,n = 1, otherwise b m,n = 0; define the beam resource allocation matrix B = [b1, b2,..., b M ] T , considering that the UAV targets in the UAV formation are relatively dispersed in space, a single beam cannot cover multiple targets at the same time, it is assumed that a beam of each interference device can only interfere with a single UAV target; due to the performance constraints of the interference device, a single interference device can simultaneously interfere with at most I UAV targets: To ensure that the jamming beam resources are reasonably utilized, the maximum number of unmanned aerial vehicles that each unmanned aerial vehicle can be jammed is limited to U: S3: Introducing the target threat degree evaluation link, constructing the interference beam resource allocation optimization problem, and guiding the reasonable allocation of interference beam resources: When allocating interference resources, the differences in the threat degrees of different UAV targets need to be considered. According to the evaluation results of the interference device m, the target threat degree evaluation vector w is defined m =[w m,1 ,w m,2 ,…,w m,N ] T , where 0≤w m,n ≤1; the evaluation results of all interference devices are integrated to define the target threat degree evaluation matrix W = [w1, w2,..., w M ] T ; and the average threat degree of the UAV target n is defined as The problem of allocating interference beams of multiple interference devices to the UAV formation can be represented as: Wherein, P1 is a jamming beam allocation problem, max represents maximization, P1 represents that the sum of the average threat degrees of all unmanned aerial vehicles is maximized by solving the beam resource allocation matrix B under the constraint conditions C1 and C2; S4: a target threat degree evaluation model is constructed based on a tactical significance map (TSM), and a target threat degree evaluation matrix is generated; S5: the TSM function is improved by using a discrete soft actor-critic (DSAC) algorithm, and the target threat degree evaluation matrix is reconstructed; S6: based on the target threat degree evaluation matrix, a jamming beam resource allocation strategy is executed, and a jamming beam resource allocation matrix is generated.

2. The deep reinforcement learning aided target threat assessment model and jamming beam allocation method according to claim 1, characterized in that: In the S1, a "many-to-many" wartime communication confrontation scenario is considered, in which M jamming devices of the interference side implement suppressive jamming on a UAV formation composed of N unmanned aerial vehicles commanded by an enemy early warning aircraft; N e ={1, 2, …, N} and M e ={1, 2, …, M} represent sets composed of N unmanned aerial vehicles in the UAV formation and M jamming devices, respectively. In order to avoid jamming, the speed and direction of the UAV formation are dynamically changed during flight to the jamming site; the jamming devices can generate multiple beams to simultaneously jam multiple UAV targets. 3.The deep reinforcement learning aided target threat assessment model and jamming beam allocation method according to claim 1, characterized in that: In the S4: assuming that the communication confrontation scene is an XY two-dimensional space, the target threat degree is evaluated based on the TSM function, and the TSM function can be expressed as: wherein, is the state vector of the UAV target n at the current moment, [x n ,y n ] T and respectively represent the position vector and the velocity vector of the UAV target n, the reference point coordinate is h c = [x c ,y c ] T , exp is an exponential function, represents the distance between the UAV target n and the reference point, k0 and m0 are preset constants, is the absolute velocity of the UAV target n, θ c,n is the heading angle of the UAV target n relative to the reference point, and its expression is arccos represents an inverse cosine function; all interference devices are taken as reference points, and target threat degree assessment is performed on the UAV target to generate a target threat degree assessment matrix W.

4. The deep reinforcement learning aided target threat assessment model and jamming beam allocation method according to claim 1, characterized in that: In the S5: when the unmanned aerial vehicle target is far away from the jamming side, the target threat degree evaluation result based on the TSM function will degenerate into the evaluation result based on a single jamming device; To solve this problem, the target threat degree evaluation matrix reconstruction problem is modeled as a Markov decision process (MDP), including a state space, an action space and a reward function; The motion state information of the UAV target n to the M interference devices at time t is s n,b,t = [o 1,b,t , o 2,b,t , ..., o M,b,t ], wherein subscript b represents a beam, subscript t represents a time, o m,b.t = [R m,n,t , v n,t , θ m,n,t ], and contains the distance R m,n,t , the speed v n,t , and the heading angle θ m,n,t of the UAV target n to the interference device m at time t; based on s n,b,t , the state of the MDP is defined as s b,t = [s 1,b,t , s 2,b,t , ..., s N,b,t ] ∈ S b , wherein S b is a state space; the action space is M e ; the reward function is based on the TSM function, and when the interference device is selected to interfere with the UAV target n, if the threat degree evaluation value w m,n of the selected interference device m is higher than all evaluation results w n = [w m,1 , w m,2 , ..., w m,N ] T , the reward value obtained is higher.

5. The deep reinforcement learning aided target threat assessment model and jamming beam allocation method according to claim 1, characterized in that: In the S5, the TSM function is improved by using the DSAC algorithm, and the reconstruction process of the target threat degree evaluation matrix includes the following steps: S51: initialize the number of interference devices M, the number of UAV targets N, the total number of episodes max_episode, the total interaction time step of each episode max_t, the parameters of the policy (Actor) network, the evaluation (Critic) network and the target evaluation (Critic) network, the experience replay pool, the threshold Φ b , let B = 0, W = 0; S52: if the number of training episodes exceeds max_episode, the training is ended, S56 is executed, otherwise max_t is initialized, and S53 is executed; S53: if the number of interaction time steps exceeds max_t, S52 is executed, otherwise S54 is executed; S54: For drone target n, the agent Actor network inputs s n,b,t , outputs action a n,b,t and obtains reward r n,b,t , the state moves to s n,b,t+1 , stores the experience sample (s n,b,t , a n,b,t , r n,b,t , s n,b,t+1 ) to the experience replay pool; S55: If the number of experience replay pool samples is greater than a threshold value Φ b , randomly sample a batch of samples, train the Actor and Critic networks, otherwise perform S53; S56: the DSAC model is saved.

6. The deep reinforcement learning aided target threat assessment model and jamming beam allocation method according to claim 1, characterized in that: In the S6, based on the target threat degree evaluation matrix, a jamming beam allocation strategy is executed, and a jamming beam resource allocation matrix is generated, which specifically includes the following steps: S61: For each UAV target n, the agent Actor network inputs s n,b,t , outputs a probability vector ξ n and saves, reconstructing the target threat assessment matrix W = [ξ1, ξ2,..., ξ N ] S62: loop through W, select the position with the maximum threat degree evaluation value in W in descending order, set it to 0, map the corresponding value in the jamming beam allocation matrix B to 1, and the mapping process needs to satisfy the constraint conditions C1 and C2 at the same time; S63: based on S62, the mapping process is repeated, and when the number of unmanned aerial vehicles allocated to each jamming device reaches the maximum number of unmanned aerial vehicle targets that can be jammed simultaneously, the beam resource allocation process is ended; S64: the jamming beam resource allocation matrix B is output.

Citation Information

Patent Citations

  • Air target sensor management method based on target threat degree

    CN110530424A

  • Unmanned aerial vehicle threat assessment method and system based on recurrent neural network

    CN116502909A