Networking radar collaborative deception jamming decision-making method based on multi-agent reinforcement learning

Through the network radar collaborative spoofing and jamming decision-making method based on multi-agent reinforcement learning, the ADM-MATD3 algorithm is used to optimize the multi-agent collaborative spoofing and jamming strategy, which solves the problems of high computing complexity and slow convergence speed in the existing technology, and achieves a more efficient network radar spoofing and jamming effect.

CN120334866APending Publication Date: 2025-07-18UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510753135.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the collaborative fraud interference decision-making framework relies on algorithms such as dynamic programming, evolutionary computing and supervised machine learning, resulting in high computational complexity, slow convergence speed and limited generalization capabilities, making it difficult to quickly generate efficient network radar fraud interference strategies.

Method used

Using a multi-agent reinforcement learning method, a network radar collaborative verification model and multi-agent collaborative fraud interference strategy optimization model are established, and the ADM-MATD3 reinforcement learning algorithm is used for strategy solving, combining artificial potential field reward function, strategic space dynamic search and multi-head self-attention mechanism to optimize the strategy of multi-agent collaborative fraud interference network radar.

Benefits of technology

It improves the efficiency of strategy search and decision-making feasibility, significantly improves the interference effect, and has faster convergence speed and higher solution efficiency than existing algorithms, which can effectively solve the problem of radar strategy optimization for multi-agent collaborative spoofing interference networking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120334866A_ABST
    Figure CN120334866A_ABST
Patent Text Reader

Abstract

The invention discloses a networking radar collaborative deception jamming decision-making method based on multi-agent reinforcement learning, and the method comprises the steps: firstly building a networking radar collaborative verification model, carrying out the verification and fusion of a networking radar, building a multi-agent collaborative deception jamming networking radar strategy optimization model, and carrying out the verification and fusion of the networking radar. A multi-agent collaborative deception jamming decision-making problem oriented to the networking radar is described as a distributed Markov decision-making process, and finally a strategy of the multi-agent collaborative deception jamming networking radar is solved by adopting an ADM-MATD3 reinforcement learning algorithm. According to the method, the intelligent agent can learn the difference between different feasible strategies in a radar beam, the strategy search efficiency is improved, the accurate strategy is output, and compared with an existing multi-intelligent-agent reinforcement learning algorithm, the method has the advantages of higher convergence speed, higher solving efficiency, higher decision feasibility and higher robustness. The problem of multi-agent collaborative deception jamming networking radar strategy optimization is effectively solved, and the jamming effect is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of radar electronic countermeasures, and particularly relates to a collaborative deception jamming decision-making method for networked radars based on multi-agent reinforcement learning. Background Art

[0002] Collaborative deception jamming technology is a widely used jamming method for networked radar detection and tracking systems. The jammer forms a series of regular false measurement points at the radar tracking end by continuously forwarding radar signals with controlled time delays in multiple frames, thereby confusing the data association of the radar tracking system, and further forcing the radar tracking system to incorrectly track false targets and lose real targets.

[0003] In the prior art, the literature "M N Wen, Reinforcing language agents via policy optimization with action decomposition. Vancouver, Canada: NeurIPS, 2024" proposed a new multi-agent policy optimization architecture, indicating that deep reinforcement learning can better handle cooperation problems in complex scenarios during the multi-objective decision-making process. Due to the advantages of agents controlled by deep reinforcement learning, such as strong adaptive generalization ability to complex scenarios, high decision-making speed, and low dependence on training samples, this intelligent decision-making paradigm based on the multi-agent policy optimization architecture provides a new technical breakthrough direction for breaking the collaborative verification and homologous fusion detection systems of networked radars.

[0004] Although many existing technologies have given corresponding collaborative decision-making schemes, the known collaborative deception jamming decision-making frameworks mainly rely on algorithm architectures such as dynamic programming, evolutionary computation, and supervised machine learning. The decision-making computational complexity of the dynamic programming framework increases exponentially with the expansion of the scenario scale, ultimately leading to difficult and slow solution; the evolutionary computation framework is prone to falling into local optima when facing complex collaborative feasible solution spaces and has a slow algorithm convergence speed; the supervised machine learning framework has severely limited model generalization ability due to the difficulty of obtaining high-quality jamming countermeasure samples in the actual combat environment. These limitations seriously restrict the generation efficiency of collaborative deception jamming networked radar strategies and the real-time deception effect. In order to cope with the severe challenges brought by the collaborative detection system to the current combat environment, it is urgent to study more efficient, intelligent, and generalized multi-agent collaborative deception jamming decision-making technologies. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a collaborative deception jamming decision-making method for networked radars based on multi-agent reinforcement learning, which solves the problem of solving multi-agent collaborative deception jamming strategies for networked radars.

[0006] The technical solution adopted by the present invention is as follows: A collaborative deception interference decision-making method for networking radars based on multi-agent reinforcement learning, and the specific steps are as follows:

[0007] S1. Establish a collaborative verification model for networking radars, and perform verification and fusion of networking radars;

[0008] S2. Based on step S1, establish a multi-agent collaborative deception interference networking radar strategy optimization model;

[0009] Combine the three-dimensional kinematic equation to establish a multi-agent motion model, then set the strong dynamic constraints suffered by the agents to obtain a multi-agent kinematic model, and transform the multi-agent collaborative control problem into a policy evaluation function optimization problem to construct a multi-agent collaborative deception interference networking radar strategy optimization model.

[0010] S3. Based on the strategy optimization model constructed in step S2, use the ADM-MATD3 reinforcement learning algorithm to solve the strategy of multi-agent collaborative deception interference networking radars;

[0011] Based on the problem of solving the strategy of multi-agent collaborative deception interference networking radars, the ADM-MATD3 reinforcement learning algorithm is proposed to optimize it. The algorithm includes three parts: the design of a reward function based on an artificial potential field, a dynamic search method for the strategy space, and a Critic neural network structure based on a multi-head self-attention mechanism. And the ADM-MATD3 reinforcement learning algorithm only introduces a multi-head self-attention mechanism in the Critic network.

[0012] Furthermore, the specific steps of step S1 are as follows:

[0013] S11. Homologous point track verification of networking radars;

[0014] First, calculate the association distance d and the association threshold G between the measurements of the same target by two different radars, and then compare the calculated association distance with the association threshold. When the association distance is less than the association threshold, it is considered that the two measurements come from the same target, and the decision rule H expression is as follows:

[0015]

[0016] Set the measurements of two different radars as M1 and M2 respectively, and then define their association distance as the Euclidean distance between the two measurements, and the expression is as follows:

[0017] d = ||M1 - M2|| (2)

[0018] Set the covariance matrices of the homologous measurements M1 and M2 as C1 and C2. For the data processing flow of networked radars, M1 and M2 are two independent measurements, that is, it can be considered that the two measurements are uncorrelated. Then, according to random analysis, the covariance matrix of the difference between the two measurements M1 - M2 is C1 + C2. After performing eigenvalue decomposition on C1 + C2, the following expression is obtained:

[0019]

[0020] where, Λ represents a diagonal matrix with the eigenvalues of C1 + C2 as diagonal elements, T represents the matrix transpose operation, W represents the vector matrix of C1 + C2, λ ∈ (λ1, λ2, λ3) represents the eigenvalues of the C1 + C2 matrix. When setting the largest eigenvalue among all eigenvalues as λ max , then the expression for defining its adaptive threshold G is as follows:

[0021]

[0022] where, G c represents the adaptive threshold coefficient. By comparing the threshold G and calculating the association distance, non-homologous traces can be eliminated.

[0023] S12, Homologous data fusion of networked radars;

[0024] Set the N homologous measurements observed by N radars as M i , i = 1, 2,..., N. Each measurement corresponds to a covariance matrix C i , i = 1, 2,..., N. Then, the fusion covariance matrix of these N homologous measurements and the fusion measurement are expressed as follows:

[0025]

[0026] where, ω i represents the weighting coefficient, which is obtained by solving through the following expression:

[0027]

[0028] Furthermore, the specific steps of step S2 are as follows:

[0029] S21, Establish a multi-agent kinematic model based on the strong kinematic constraints suffered by the agents;

[0030] The instantaneous expression of the agent kinematic model is as follows:

[0031]

[0032] where, Δx t , Δy t, Δz t respectively represent the motion change amounts of a single agent on the x, y, and z axes within Δt time in the ground absolute coordinate system; v t , θ t respectively represent the velocity, azimuth angle, and pitch angle of a single agent at time t in the ground absolute coordinate system.

[0033] The constraint expressions for the agent's velocity, azimuth angle, and pitch angle are as follows:

[0034]

[0035] Among them, v min , v max respectively represent the minimum velocity and maximum velocity of the agent; θ min , θ max respectively represent the minimum pitch angle and maximum pitch angle of the agent; respectively represent the minimum azimuth angle and maximum azimuth angle of the agent; a t represents the acceleration decided by the agent at time t; represents the azimuth angle change amount within Δt time in the ground absolute coordinate system decided by the agent at time t, Δθ t represents the pitch angle change amount within Δt time in the ground absolute coordinate system decided by the agent at time t, a max represents the maximum acceleration of the agent in the ground absolute coordinate system, Δθ max represents the maximum pitch angle change amount of the agent in the ground absolute coordinate system, represents the maximum azimuth angle change amount of the agent in the ground absolute coordinate system.

[0036] S22. Based on step S21, construct an optimization model for the multi-agent cooperative deception jamming networking radar strategy;

[0037] Describe the coordinate change of the agent at time t as the influence brought by the control strategy The expression is as follows:

[0038]

[0039] Among them, respectively represent the distance change amounts of each axis of agent m at time t in the three-dimensional coordinate system.

[0040] Finally, transform the multi-agent cooperative control problem into a policy evaluation function optimization problem, and its optimization model expression is as follows:

[0041]

[0042] Among them, M represents the number of agents, Γ represents the feasible solution space of the agent kinematic constraints, and [ξ1(t), …, ξ M (t)] represents the strategies selected by each agent at time t, represents the entire time series and is the evaluation function of the overall effectiveness of all agents' strategies within it.

[0043] Furthermore, the specific steps of step S3 are as follows:

[0044] S31. Design a reinforcement learning reward function based on the artificial potential field APF;

[0045] First, the basic reward is set as the negative of the distance from the agent to the LOS area, so that the agent is attracted by the LOS area at all times. Then, combining the method of 0 / 1 reward, the area is divided into two parts: inside and outside the radar beam. The reward for the agent being outside the radar beam is -1, and the reward for being inside the radar beam is 1. Then, the corresponding reward function expressions can be obtained inside / outside the radar beam as follows:

[0046]

[0047] Among them, represents the reward obtained by agent m at the t-th moment of decision-making, represents the perpendicular distance from agent m to the line connecting the towing false point and the radar coordinates at the t-th moment, represents the actual width of the beam corresponding to the perpendicular point of agent m at the t-th moment.

[0048] Then, based on the radar beam width, the reward function is further corrected, and the expression is as follows:

[0049]

[0050] Among them, S(·) represents the scaling function, λ B represents the beam discount coefficient, and B min represents the minimum allowable discounted beam radius.

[0051] S32. Policy space dynamic search method;

[0052] The exploration rate changes with the training steps, and the expression of the exploration rate std(x) is as follows:

[0053] std(x) = softmax(f(x)) x = 1, 2, …, K (14)

[0054] Among them, K represents the maximum number of training steps per episode, softmax(·) represents the normalized exponential function, and the expression of f(x) is as follows:

[0055]

[0056] Among them, C represents the current number of trainings, E max represents the maximum number of trainings, and ST represents the number of times the sliding window slides during training.

[0057] S33. The Critic neural network architecture based on the multi-head attention mechanism;

[0058] Define the query matrix Q, the key matrix K, the value matrix V, the normalization exponential function softmax(·), and the dimension dim of the key matrix K , then the expression of single-head scaled dot-product attention is as follows:

[0059]

[0060] The multi-head self-attention mechanism is improved based on self-attention. The self-attention is copied N times, and each copy has its own set of queries, keys, and values. After calculating the self-attention values of these N copies respectively, they are concatenated and output to obtain the multi-head self-attention value.

[0061] For the i-th head, its corresponding query, key, and value are the results after the original query, key, and value are projected into the i-th subspace through different linear transformations. The expression is as follows:

[0062]

[0063] Among them, W i Q , W i K , W i V respectively represent the learnable weight matrices corresponding to the query Q, key K, and value V of the i-th head. This matrix is continuously updated during training.

[0064] In the multi-head attention mechanism, the calculation expression of each head attention is as follows:

[0065] head i =Attention(Q i , K i , V i ) (18)

[0066] After calculating all single-head attentions, the results of each head attention are concatenated, and then multiplied by the weights of each head to output the result of the multi-head attention mechanism. The expression is as follows:

[0067] MHA(Q, K, V) = Concat(head1, head2, …, head N )W O (19)

[0068] Among them, W O represents the output weight of the multi-head self-attention, and head i represents the i-th head.

[0069] S34. Based on steps S31 - S33, use the ADM-MATD3 reinforcement learning algorithm to solve the strategy of multi-agent collaborative deception jamming of the networking radar;

[0070] S341. Initialize the coordinates of the multi-agent cluster and the coordinates of the networking radar, set the number of randomly sampled experiences for each single training of the reinforcement learning, initialize the number of delayed update rounds and the maximum number of training rounds;

[0071] S342. Initialize the value network, policy network parameters of all agents and the experience replay pool, and reset the environment to obtain the initial state s0;

[0072] S343. Each agent independently observes the environment to obtain the current state

[0073] S344. After each agent inputs its own state into its respective policy network, it outputs its own action a, and a is superimposed with perturbations according to the dynamic search rate:

[0074] Among them, represents the output obtained after the current state is input into the policy network, ε represents the Gaussian noise that conforms to the dynamic exploration variance of the current round, and std(episode) represents the dynamic exploration variance of the current round calculated in formula (14).

[0075] S345. The agent executes the action a, and the environment performs a state transition according to the agent's action, outputs the reward and the next moment state s′, and stores the obtained experience in the experience pool;

[0076] S346. If the delayed update condition is met, go to step S347; otherwise, jump back to step S343;

[0077] Among them, the delayed update means that after every preset number of update rounds, the agent will perform neural network training and gradient descent update, and during the training rounds, the agent only conducts exploration and experience accumulation.

[0078] S347. Based on the preset number of randomly sampled experience values in step S341, extract a certain amount of experience from the experience pool to train the value network and the policy network;

[0079] S348. Use the gradient descent method to update the policy network and the soft update method to update the target network;

[0080] S349. If the number of training rounds is equal to the set maximum number of training rounds, save the ADM-MATD3 network model; otherwise, reset the environment, and loop through steps S344 - S349 until the set number of training rounds is reached.

[0081] Beneficial effects of the present invention: The method of the present invention first establishes a collaborative verification model for networked radars and conducts verification and fusion of networked radars. Then, it establishes a multi-agent collaborative deception interference networked radar strategy optimization model, describes the multi-agent collaborative deception interference decision-making problem for networked radars as a distributed Markov decision process, and finally uses the ADM-MATD3 reinforcement learning algorithm to solve the strategy of multi-agent collaborative deception interference for networked radars. The method of the present invention enables agents to learn the differences between different feasible strategies within the radar beam while improving the strategy search efficiency and outputting accurate strategies. Compared with existing multi-agent reinforcement learning algorithms, it has a faster convergence speed, higher solution efficiency, and decision feasibility, effectively solving the problem of multi-agent collaborative deception interference networked radar strategy optimization and significantly improving the interference effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 It is a flowchart of a method for collaborative deception interference decision-making of networked radars based on multi-agent reinforcement learning according to the present invention.

[0083] Figure 2 It is a schematic diagram of the principle of homologous point track inspection for networked radars in an embodiment of the present invention.

[0084] Figure 3 It is a final effect demonstration diagram of fixed-wing UAVs collaborating to deceive and interfere with networked radars in an embodiment of the present invention.

[0085] Figure 4 It is a schematic diagram of the specific numerical simulation results of the dynamic exploration rate in an embodiment of the present invention.

[0086] Figure 5 It is a diagram of the flight range change of UAVs at the dynamic exploration rate in an embodiment of the present invention.

[0087] Figure 6 It is an architecture diagram of the Critic network based on the multi-head self-attention mechanism in an embodiment of the present invention.

[0088] Figure 7 It is a schematic diagram of the performance curve comparison between the ADM-MATD3 algorithm and the MATD3 algorithm trained with the same parameters in an embodiment of the present invention.

[0089] Figure 8 It is a schematic diagram of the UAV flight trajectory and towing point track obtained by the ADM-MATD3 decision model in an embodiment of the present invention.

[0090] Figure 9This is a schematic diagram of the collaborative towing point track result after the networking radar fusion in the embodiment of the present invention. Detailed implementation manners

[0091] The method of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0092] To facilitate the description of the content of the present invention, the following terms are first explained:

[0093] Term 1: Collaborative deception jamming

[0094] Multiple agents capable of executing deception jamming strategies respectively perform range deception jamming on different radars in the networking system, and thus false point tracks can be formed in the monitoring window of the networking radar.

[0095] Term 2: MARL

[0096] MARL is the English abbreviation of MultiAgent Reinforcement Learning, which is a machine learning method aiming to enable agents to learn how to make optimal decisions in specific tasks by interacting with the environment.

[0097] Term 3: MATD3

[0098] MATD3 is the English abbreviation of the Multi Agent Twin Delayed DeepDeterministic Policy Gradient algorithm, which improves the stability of the algorithm through a dual Q-network and a delayed update strategy.

[0099] As Figure 1 shown, the flowchart of a networking radar collaborative deception jamming decision method based on multi-agent reinforcement learning of the present invention is as follows:

[0100] S1. Establish a networking radar collaborative verification model, and perform networking radar verification and fusion;

[0101] S2. Based on step S1, establish a multi-agent collaborative deception jamming networking radar strategy optimization model;

[0102] Establish a multi-agent motion model in combination with the three-dimensional kinematic equation, then set the strong dynamic constraints suffered by the agents to obtain a multi-agent kinematic model, and construct a strategy optimization model for multi-agent collaborative deception jamming of the networking radar by transforming the multi-agent collaborative control problem into a strategy evaluation function optimization problem.

[0103] S3. Based on the policy optimization model constructed in step S2, the ADM-MATD3 reinforcement learning algorithm is used to solve the policy of multi-agent collaborative deception jamming of a netted radar;

[0104] Regarding the problem of solving the policy of multi-agent collaborative deception jamming of a netted radar, the ADM-MATD3 reinforcement learning algorithm is proposed for optimization. The algorithm includes three parts: the design of a reward function based on artificial potential field, a dynamic search method for the policy space, and a Critic neural network structure based on the multi-head self-attention mechanism.

[0105] The reward function based on artificial potential field is used to provide return values to precisely control the agents. The dynamic search method for the policy space can improve the search efficiency for high-dimensional feasible solutions, while the Critic neural network structure based on the multi-head self-attention mechanism can focus on the association between the input and return values from different perspectives to accelerate the algorithm's convergence speed.

[0106] In the ADM-MATD3 reinforcement learning algorithm, the behavior of the Actor network is restricted and corrected by the Critic network. Therefore, only by introducing the multi-head self-attention mechanism into the Critic network can the algorithm's ability to focus on details be enhanced.

[0107] In this embodiment, the agent uses a fixed-wing unmanned aerial vehicle. Due to the advantages of fixed-wing unmanned aerial vehicles such as flexible cooperative combat methods, strong concealment capabilities, large upper limits of flight speed, and long maximum ranges, how to efficiently and robustly control a group of fixed-wing unmanned aerial vehicles to collaboratively deceive and jam a netted radar has become the key to completing tasks such as covering the penetration of high-value platforms and deceiving radar detection in the netted radar detection system.

[0108] In this embodiment, step S1 is specifically as follows:

[0109] S11. Homologous point track verification of the netted radar;

[0110] When a target is observed by multiple radars simultaneously, each radar obtains a measurement with different coordinates but the same origin. The processing center aggregates all measurement data and then performs homologous point track verification on the measurement data to eliminate point tracks from different sources to achieve anti-deception jamming.

[0111] First, calculate the association distance d between the measurements of the same target by two different radars and the association threshold G, and then compare the calculated association distance with the association threshold. When the association distance is less than the association threshold, it is considered that the two measurements come from the same target. The decision rule H expression is as follows:

[0112]

[0113] Set two different radar measurements as M1 and M2 respectively, and then define their association distance as the Euclidean distance between the two measurements. The expression is as follows:

[0114] d = ||M1 - M2|| (2)

[0115] Set the covariance matrices of the homologous measurements M1 and M2 as C1 and C2. For the data processing flow of the networking radar, M1 and M2 are two independent measurements, that is, it can be considered that the two measurements are uncorrelated. Then, the covariance matrix of the difference between the two measurements M1 - M2 is obtained as C1 + C2 according to random analysis. After performing eigenvalue decomposition on C1 + C2, the expression is as follows:

[0116]

[0117] Among them, Λ represents a diagonal matrix with the eigenvalues of C1 + C2 as diagonal elements, T represents the matrix transpose operation, W represents the vector matrix of C1 + C2, λ ∈ (λ1, λ2, λ3) represents the eigenvalues of the C1 + C2 matrix. When setting the largest eigenvalue among all eigenvalues as λ max , then the expression for defining its adaptive threshold G is as follows:

[0118]

[0119] Among them, G c represents the adaptive threshold coefficient. In the project of this embodiment, when the radar signal reception probability is 99.74%, G c takes 3. By comparing the threshold G and calculating the association distance, non-homologous traces can be eliminated. The specific method is as Figure 2 shown.

[0120] S12. Homologous data fusion of the networking radar;

[0121] After multiple radars obtain all homologous measurements that pass the homologous trace verification, it is necessary to perform data fusion on these multiple homologous measurements to determine the target position as accurately as possible. The essence of homologous data fusion is the weighted summation of multiple homologous measurements.

[0122] Set N radars to observe N homologous measurements as M i , i = 1, 2,..., N, and the covariance matrix corresponding to each measurement is C i , i = 1, 2,..., N. Then, the fusion covariance matrix of these N homologous measurements and the fusion measurement are expressed as follows:

[0123]

[0124] Among them, ω i represents the weighting coefficient, which is obtained by solving through the following expression:

[0125]

[0126] In this embodiment, step S2 is specifically as follows:

[0127] S21. Establish a multi-agent kinematic model based on the strong kinematic constraints suffered by the agents;

[0128] In this embodiment, the agent is a fixed-wing UAV. The aerodynamic limitations suffered by the fixed-wing UAV result in poor flexibility, and stronger kinematic constraints are imposed during the execution of flight missions. Then, the instantaneous expression of its kinematic model is as follows:

[0129]

[0130] where, Δx t , Δy t , Δz t respectively represent the motion changes of a single agent on the x, y, and z axes within Δt time in the ground absolute coordinate system; v t , θ t respectively represent the speed, azimuth angle, and pitch angle of a single agent at time t in the ground absolute coordinate system.

[0131] The constraint expressions for the agent's speed, azimuth angle, and pitch angle are as follows:

[0132]

[0133] where, v min , v max respectively represent the minimum speed and maximum speed of the agent; θ min , θ max respectively represent the minimum pitch angle and maximum pitch angle of the agent; respectively represent the minimum azimuth angle and maximum azimuth angle of the agent; a t represents the acceleration decision made by the agent at time t; represents the azimuth angle change amount within Δt time in the ground absolute coordinate system decided by the agent at time t, Δθ t represents the pitch angle change amount within Δt time in the ground absolute coordinate system decided by the agent at time t, a max represents the maximum acceleration of the agent in the ground absolute coordinate system, Δθ max represents the maximum pitch angle change amount of the agent in the ground absolute coordinate system, represents the maximum azimuth angle change amount of the agent in the ground absolute coordinate system.

[0134] S22. Based on step S21, construct an optimization model for the multi-agent collaborative deception jamming networking radar strategy;

[0135] Describe the change in the agent's coordinates at time t as the impact brought about by the control strategy The expression is as follows:

[0136]

[0137] Among them, respectively represent the change in distance of each axis of agent m at the t-th moment in the three-dimensional coordinate system.

[0138] Finally, transform the multi-agent cooperative control problem into a policy evaluation function optimization problem, and its optimization model expression is as follows:

[0139]

[0140] Among them, M represents the number of agents, Γ represents the feasible solution space of the agent kinematic constraints, and [ξ1(t),…,ξ M (t)] represents the strategies selected by each agent at time t, represents the entire time series is the evaluation function for the overall effectiveness of all agents' strategies within the time series, used to evaluate the comprehensive effectiveness of the strategies obtained from multiple decisions, Figure 3 which is the final effect demonstration diagram of the cooperative deception interference networking radar of the fixed-wing UAV in this embodiment.

[0141] In this embodiment, the specific steps of step S3 are as follows:

[0142] S31. Design a reinforcement learning reward function based on the artificial potential field APF (Artificial Potential Field);

[0143] First, the basic reward is set as the negative of the distance from the agent to the LOS area, so that the agent is attracted by the LOS area at all times. Then, in view of the problem that the distance change gradient is relatively gentle and the gravitational force is relatively weak, the area is divided into two parts, inside and outside the radar beam, by combining the 0 / 1 reward method. The reward for the agent being outside the radar beam is -1, and the reward for being inside the radar beam is 1, to accelerate the attraction of the agent to fly into the radar beam. The corresponding reward function expressions can be obtained inside / outside the radar beam as follows:

[0144]

[0145] Among them, represents the reward obtained by agent m at the t-th moment of decision-making, represents the perpendicular distance from agent m at the t-th moment to the line connecting the towing false point and the radar coordinates, represents the actual width of the beam corresponding to the perpendicular point of agent m at the t-th moment.

[0146] In the process of combining the adversarial environment with reinforcement learning within the LOS region, another problem will be faced: the environmental size results in an excessive order of magnitude of the distance parameter. Since it is not recommended to use too large a value as a reward in reinforcement learning, when distance is introduced as a reference in the reward function, methods such as normalization or proportional reduction are adopted to use the distance value as a reward. Using the normalized distance as the reward function makes it difficult for deep reinforcement learning to provide an accurate numerical solution during function approximation. Therefore, the reward function described in Equation (12) in this embodiment can only guide the UAV to the LOS region, and the part closer to the LOS region is difficult to further refine.

[0147] Based on this, the reward function is further corrected within the radar beam width, and the expression is as follows:

[0148]

[0149] where S(·) represents the scaling function, and λ B represents the beam discount coefficient. In this embodiment, λ B realizes that as the training progresses, the gradient of the reward function gradually approaches the beam axis to better guide the UAV by gradually reducing the beam width considered by the UAV during the training process. B min represents the minimum allowable discounted beam radius.

[0150] S32. Policy space dynamic search method;

[0151] The problem of the exponential explosion of the agent's policy search space in the sequential decision-making process is solved through the policy space dynamic search method. The exploration rate changes with the training steps. As Figure 4 shown, the expression of the exploration rate std(x) is as follows:

[0152] std(x) = softmax(f(x)) x = 1, 2, …, K (14)

[0153] where K represents the maximum number of training steps per episode, softmax(·) represents the normalized exponential function, whose role is to smooth the exploration rate that changes with each episode, so that its change amount will not be too high. The expression of f(x) is as follows:

[0154]

[0155] where C represents the current training count, and E max represents the maximum training count, and ST represents the number of sliding times of the sliding window during training. In this embodiment, the size of the UAV policy space obtained through Equation (15) is as Figure 5 shown (the change in the flight range of the UAV under the dynamic exploration rate). The exploration rate of the action output by the Actor network slowly changes as the number of training episodes increases. The specific benefits are as follows:

[0156] (1) In the initial training stage, increasing the exploration rate in the early stage and decreasing the exploration rate in the later stage within the same episode of the agent can enable the network to quickly explore better strategies or optimal feasible trends in the early stage, laying a decision-making foundation for exploration in the middle and later stages of the same episode;

[0157] (2) Selecting a small and almost stable exploration rate in the middle training stage can stabilize the exploration results of the Actor in the initial stage;

[0158] (3) In the final training stage, the exploration rate of the agent becomes small in the early stage and large in the later stage, and its effect is to better search for later-stage strategies based on the basically unchanged early-stage strategies.

[0159] In summary, the strategy space search method based on the dynamic exploration rate can effectively improve the search speed of feasible strategies and reduce the number of training episodes required for the algorithm to converge.

[0160] S33. The Critic neural network architecture based on the multi-head attention mechanism;

[0161] Define the query matrix Q, the key matrix K, the value matrix V, the normalization exponential function softmax(·), and the dimension dim of the key matrix K , then the expression of the single-head scaled dot-product attention is as follows:

[0162]

[0163] Equation (16) introduces the scaling factor dim K , and its purpose is to avoid the problem of gradient vanishing or gradient explosion caused by the too large dimension of the key matrix when calculating the dot product, so as to ensure the training stability.

[0164] The multi-head self-attention mechanism is improved based on self-attention. The self-attention is copied N times, and each copy has its own set of query, key, and value. After calculating the self-attention values of these N copies respectively, they are concatenated and output to obtain the multi-head self-attention value. This approach allows the neural network model to evaluate the input from multiple perspectives respectively.

[0165] For the i-th head, its corresponding query, key, and value are the results after the original query, key, and value are projected into the i-th subspace through different linear transformations, and the expressions are as follows:

[0166]

[0167] Among them, W i Q , W i K , W i VRespectively represent the learnable weight matrices corresponding to the query Q, key K, and value V for the i-th head. This matrix is continuously updated during training.

[0168] In the multi-head attention mechanism, the calculation expression for each head attention is as follows:

[0169] head i = Attention(Q i , K i , V i ) (18)

[0170] After calculating all the single-head attentions, the results of each head attention are concatenated, and then multiplied by the weight of each head to output the result of the multi-head attention mechanism. The expression is as follows:

[0171] MHA(Q, K, V) = Concat(head1, head2,..., head N )W O (19)

[0172] Among them, W O represents the output weight of the multi-head self-attention, and head i represents the i-th head. Figure 6 This is the architecture diagram of the Critic network in this embodiment. In the figure, a M*1 represents the actions selected by each agent, O M*obs_dim represents the observations of each agent, and the size of each observation is obs_dim. Q M*k , K M*k , V M*k respectively correspond to the query, key, and value of each agent output by the linear layer. Q M*1 represents the Q value of each agent's reinforcement learning calculated by the multi-head self-attention mechanism.

[0173] For the MATD3 algorithm based on the AC framework, the behavior of the Actor network is restricted and corrected by the Critic network. Therefore, only by introducing the multi-head self-attention mechanism into the Critic network can the algorithm's ability to focus on details be enhanced, and the Actor network still uses the existing deep reinforcement learning network architecture.

[0174] S34. Based on steps S31 - S33, use the ADM-MATD3 reinforcement learning algorithm to solve the strategy of the fixed-wing UAV cooperative deception jamming networking radar;

[0175] S341. Initialize the coordinates of the multi-agent cluster and the coordinates of the networking radar, set the number of experiences randomly sampled in a single training of the reinforcement learning, initialize the number of delayed update rounds and the maximum number of training rounds;

[0176] S342. Initialize the value network, policy network parameters of all agents and the experience replay pool, and reset the environment to obtain the initial state s0;

[0177] S343. Each agent independently observes the environment to obtain the current state

[0178] S344. After each agent inputs its own state into its respective policy network, it outputs its own action a, and a is superimposed with disturbance according to the dynamic search rate:

[0179] where, represents the output obtained after the current state is input into the policy network, ε represents Gaussian noise that conforms to the dynamic exploration variance of the current round, and std(episode) represents the dynamic exploration variance of the current round calculated in Equation (14).

[0180] S345. The agent executes the action a, and the environment performs a state transition according to the agent's action, outputs a reward and the next moment state s′, and stores the obtained experience in the experience pool;

[0181] S346. If the delayed update condition is satisfied, go to step S347, otherwise jump back to step S343;

[0182] Among them, the delayed update means that after every preset number of update rounds, the agent will perform neural network training and gradient descent update. During the training rounds, the agent only performs exploration and experience accumulation. This approach can alleviate the overfitting of the neural network while fully searching the current policy.

[0183] S347. Based on the preset random extraction of the number of experience values in step S341, extract a fixed amount of experience from the experience pool to train the value network and the policy network;

[0184] S348. Update the policy network using the gradient descent method and update the target network using the soft update method;

[0185] S349. If the number of training rounds is equal to the set maximum number of training rounds, save the ADM-MATD3 network model, otherwise reset the environment, and loop steps S344 - S349 until the set number of training rounds is reached.

[0186] This embodiment further verifies the collaborative decision-making effect of the method of the present invention through simulation experiments. All steps and conclusions of model training and testing are carried out on the Python 3.10 platform based on the Pytorch 2.1.0 framework, and the remaining steps are verified on Matlab 2016b. Specifically as follows:

[0187] Define the proportion of the number of frames successfully passing the homologous track verification in the total preset number of frames as the track integrity, and use the final calculation result of DTW (Dynamic Time Warping) to evaluate the track similarity.

[0188] DTW can effectively calculate the similarity between two independent time series with time offset or inconsistent length, and is especially suitable for dealing with the problem of non-linear time series matching. Therefore, in this embodiment, the DTW value is used as the standard to evaluate the similarity between the false tracks after fusion and the expected preset tracks, thus providing a more reliable basis for the performance evaluation of the ADM-MATD3 algorithm.

[0189] Suppose there are two multi-dimensional time series X and Y, and the expressions are as follows:

[0190]

[0191] Among them, D represents the comparison sequence dimension, and I and J represent the lengths of the two sequences respectively.

[0192] To calculate the DTW distance between two independent sequences, it is necessary to first obtain the local distance matrix D of sequences X and Y I×J , and each element in this matrix represents the local distance between two points of the two sequences. Let X i =(x 1i , x 2i ,…, x Di )′ represent the i-th column of sequence X, and Y j =(y 1i , y 2i ,…, y Di )′ represent the j-th column of sequence Y. Then the local distance d ED (X i , Y i ) is expressed as follows:

[0193]

[0194] Equation (21) is the local distance calculated based on the Euclidean distance. The Euclidean distance assumes that all local distances have the same weight, which is not applicable in this embodiment. Therefore, the Mahalanobis distance is used to calculate the local distance between the two sequences.

[0195] Based on the Mahalanobis distance, the local distance between sequences X and Y is expressed as follows:

[0196]

[0197] Among them, represents the Mahalanobis matrix, which is a symmetric positive semi-definite matrix.

[0198] The DTW distance between the multi-dimensional sequences X and Y is calculated using the dynamic programming algorithm, that is, by combining the DTW based on the Mahalanobis matrix with dynamic programming. The expression is as follows:

[0199]

[0200] Among them, r i,j represents the local distance on the path from the first point X I×J of sequence X to the first point Y 0 of sequence Y in the distance matrix D 0 accumulating step by step to the i-th point X i of sequence X and the j-th point Y j of sequence Y. represents the similarity between sequences X and Y. In this embodiment, the above algorithm is used to compare the towing effect after convergence. The smaller the value, the higher the similarity.

[0201] It should be noted that the distance calculated by DTW is a unitless value and has no further practical significance without a reference object. Therefore, in this embodiment, the size of the DTW values of the complete algorithm and the ablation algorithm is compared to characterize the quality of the algorithm and prove the effectiveness of the algorithm.

[0202] The specific simulation scenario of this embodiment is as follows:

[0203] To evaluate the fixed-wing UAV cooperative deception jamming networking radar strategy solving algorithm proposed in this embodiment, a scenario of 4 fixed-wing UAVs against 2 networking radars is set to train the algorithm. The specific parameters of the training scenario are shown in Table 1, and the strong dynamic constraints on the UAVs are shown in Table 2 (UAV motion parameters).

[0204] Table 1

[0205] Parameter Value Number of UAVs / unit 4 Number of radars / unit 2 Number of preset false tracks / unit 2 Radar beam width / rad 0.05236 Maximum radar detection range / km 400 Radar frame period / s 7.2

[0206] Table 2

[0207]

[0208] In this embodiment, the ADM-MATD3 algorithm and its MATD3 algorithm are trained using the same parameters, and the parameters are shown in Table 3 (neural network training parameters).

[0209] Table 3

[0210] Parameter Value Maximum number of training rounds 300k Memory pool size 40000 Reward decay factor γ 0.99 Size of single training batchsize 256 Learning rate of Actor network 1e-4 Learning rate of Critic network 5e-5 Soft update parameter τ 0.1 Frequency of training performance evaluation 1000

[0211] After the ADM-MATD3 algorithm optimizes the fixed-wing UAV cooperative deception jamming networking radar strategy, a feasible UAV flight trajectory can finally be obtained, and the point trace implementation integrity and point trace fusion similarity are shown in Table 4.

[0212] Table 4

[0213]

[0214] The method of the present invention is compared with the MATD3 algorithm: According to Figure 7 the analysis of the experimental results shown, the MATD3 algorithm shows obvious instability when solving the set complex scenario, its convergence performance is poor, and it is difficult to reach a stable state within a reasonable time. Therefore, it is difficult to effectively output a reliable cooperative deception control strategy. In contrast, the proposed ADM-MATD3 algorithm of the method of the present invention shows significant advantages in the same scenario. It can not only converge quickly, but also maintain high stability during the convergence process, so as to generate an effective cooperative control strategy more efficiently. According to Figure 8 the visualization results of the strategy, it can be seen that the UAVs fly smoothly and the cooperative deception towing traces fit well, while Figure 9 the experimental results shown prove that the cooperative control strategy obtained by the ADM-MATD3 algorithm can effectively pass the homologous trace verification of the networked radar to deceive and interfere with the networked radar. It can be seen that in the complex high-dimensional cooperative decision-making problem, the ADM-MATD3 algorithm has stronger adaptability and reliability than the original MATD3 algorithm.

[0215] In summary, the method of the present invention overcomes the problems of poor single-machine accuracy and slow algorithm convergence during the cooperative deception interference of agents. It models the process of multi-agent cooperative deception interference with networked radars as a distributed Markov decision process, designs a non-linear multi-valued reinforcement learning reward function, and proposes the ADM-MATD3 algorithm to solve the strategy of multi-agent cooperative deception interference with networked radars. This algorithm can enable agents to learn the differences between different feasible strategies within the radar beam while improving the strategy search efficiency, and has a faster convergence speed compared with existing multi-agent reinforcement learning algorithms. The method of the present invention effectively solves the problem of optimizing the strategy of multi-agent cooperative deception interference with networked radars and significantly improves the interference effect.

[0216] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations without departing from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A collaborative deception jamming decision-making method for networked radars based on multi-agent reinforcement learning, the specific steps are as follows: S1. Establish a collaborative verification model for networked radars, and perform verification and fusion of networked radars; S2. Based on step S1, establish a multi-agent collaborative deception jamming networked radar strategy optimization model; Combine the three-dimensional kinematic equation to establish a multi-agent motion model, then set the strong dynamic constraints on the agents to obtain a multi-agent kinematic model, and transform the multi-agent collaborative control problem into a policy evaluation function optimization problem to construct a multi-agent collaborative deception jamming networked radar strategy optimization model; S3. Based on the strategy optimization model constructed in step S2, use the ADM-MATD3 reinforcement learning algorithm to solve the strategy of multi-agent collaborative deception jamming networked radars; Based on the problem of solving the multi-agent collaborative deception jamming networked radar strategy, propose the ADM-MATD3 reinforcement learning algorithm to optimize it; the algorithm includes three parts: the design of the reward function based on the artificial potential field, the dynamic search method of the policy space, and the Critic neural network structure based on the multi-head self-attention mechanism; and the ADM-MATD3 reinforcement learning algorithm only introduces the multi-head self-attention mechanism into the Critic network.

2. The collaborative deception jamming decision-making method for networking radars based on multi-agent reinforcement learning according to claim 1, characterized in that The specific steps of step S1 are as follows: S11. Verification of homologous tracks of networked radars; First, calculate the association distance d and the association threshold G of the measurements of the same target by two different radars, and then compare the calculated association distance with the association threshold. When the association distance is less than the association threshold, it is considered that the two measurements come from the same target. The decision rule H expression is as follows: Set the measurements of two different radars as M1 and M2, and then define their association distance as the Euclidean distance between the two measurements. The expression is as follows: d = ||M1 - M2|| (2) Set the covariance matrices of the homologous measurements M1 and M2 as C1 and C2. For the data processing flow of networked radars, M1 and M2 are two independent measurements, that is, it can be considered that the two measurements are uncorrelated. Then the covariance matrix of the difference between the two measurements M1 - M2 is obtained as C1 + C2 according to random analysis. After performing eigenvalue decomposition on C1 + C2, the expression is as follows: where, Λ represents a diagonal matrix with diagonal elements being the eigenvalues of C1 + C2, T represents the matrix transpose operation, W represents the C1 + C2 vector matrix, λ ∈ (λ1, λ2, λ3) represents the eigenvalues of the C1 + C2 matrix, and when setting the largest eigenvalue among all eigenvalues to be λ max , then the defining expression of its adaptive threshold G is as follows: Among them, G c represents the adaptive threshold coefficient. By comparing the threshold G and calculating the correlation distance, non-homologous traces can be eliminated; S12. Fusion of homologous data of networked radars; Suppose that N radars observe N homologous measurements, denoted as M i , where i = 1, 2, …, N, and the covariance matrix corresponding to each measurement is C i , where i = 1, 2, …, N. Then, the fused covariance matrix of these N homologous measurements and the fused measurement are expressed as follows: where ω i represents a weighting coefficient, which is obtained by solving the following expression:

3. A collaborative deception jamming decision-making method for networking radars based on multi-agent reinforcement learning according to claim 1, characterized in that, The specific steps of step S2 are as follows: S21. Establish a multi-agent kinematic model based on the strong kinematic constraints on the agents; The instantaneous expression of the agent kinematic model is as follows: where, Δx t , Δy t , Δz t respectively represent the motion change amounts of a single agent on the x, y, and z axes within a time interval of Δt in the ground absolute coordinate system; v t , θ t respectively represent the velocity, azimuth angle, and pitch angle of a single agent at time t in the ground absolute coordinate system; The constraint expressions of the agent speed, azimuth angle, and elevation angle are as follows: Among them, v min and v max represent the minimum speed and the maximum speed of the agent respectively; θ min and θ max represent the minimum pitch angle and the maximum pitch angle of the agent respectively; represent the minimum azimuth angle and the maximum azimuth angle of the agent respectively; a t represents the acceleration decided by the agent at time t; represents the change in azimuth angle within Δt time in the ground absolute coordinate system decided by the agent at time t, Δθ t represents the change in pitch angle within Δt time in the ground absolute coordinate system decided by the agent at time t, a max represents the maximum acceleration of the agent in the ground absolute coordinate system, Δθ max represents the maximum change in pitch angle of the agent in the ground absolute coordinate system, represents the maximum change in azimuth angle of the agent in the ground absolute coordinate system; S22. Based on step S21, construct a multi-agent collaborative deception jamming networked radar strategy optimization model; Describe the change in the agent's coordinates at time t as the impact P brought about by the control strategy t J,m , and the expression is as follows: Among them, respectively represent the distance change amounts of each axis of the agent m at the t-th moment in the three-dimensional coordinate system; Finally, transform the multi-agent collaborative control problem into a policy evaluation function optimization problem. The expression of its optimization model is as follows: Among them, \(M\) represents the number of agents, \(\Gamma\) represents the feasible solution space of the agent kinematic constraints, and \([\xi_1(t),\ldots,\xi M (t)]\) represents the strategies selected by each agent at time \(t\). represents the entire time series The evaluation function of the overall effectiveness of all agent strategies within.

4. A collaborative deception interference decision-making method for networking radars based on multi-agent reinforcement learning according to claim 1, characterized in that The specific steps of step S3 are as follows: S31. Design a reinforcement learning reward function based on the artificial potential field APF; First, the basic reward is set as the negative value of the distance from the agent to the LOS area, so that the agent is always attracted by the LOS area; then, combined with the 0 / 1 reward method, the area is divided into two parts: inside and outside the radar beam. The reward for the agent outside the radar beam is -1, and the reward for the agent inside the radar beam is 1. Then, the corresponding reward function expressions can be obtained inside / outside the radar beam as follows: Among them, represents the reward obtained by the agent m at the t-th moment of decision-making, represents the perpendicular distance from the agent m at the t-th moment to the line connecting the towed false point and the radar coordinate, represents the actual beam width corresponding to the perpendicular point of the agent m at the t-th moment; Based on this, the reward function is further corrected within the radar beam width, and the expression is as follows: Among them, S(·) represents the scaling function, and λ B represents the beam discount coefficient, and B min represents the minimum allowable discounted beam radius; S32. Policy space dynamic search method; The exploration rate changes with the training steps, and the expression of the exploration rate std(x) is as follows: std(x) = softmax(f(x)) x = 1, 2, …, K (14) where K represents the maximum number of training steps in each episode, softmax(·) represents the normalized exponential function, and the expression of f(x) is as follows: Among them, C represents the current number of trainings, and E max represents the maximum number of trainings, and ST represents the number of times the sliding window slides during training; S33. Critic neural network architecture based on the multi-head attention mechanism; Define the query matrix Q, the key matrix K, the value matrix V, the normalization exponential function softmax(`), and the dimension dim of the key matrix K , then the expression of single-head scaled dot-product attention is as follows: The multi-head self-attention mechanism is an improvement based on self-attention. The self-attention is copied N times, and each copy has its own set of query, key, and value. After calculating the self-attention values of these N copies respectively, they are concatenated and output to obtain the multi-head self-attention value; For the i-th head, its corresponding query, key, and value are the results after the original query, key, and value are projected into the i-th subspace through different linear transformations. The expression is as follows: Among them, W i Q , W i K , W i V respectively represent the learnable weight matrices corresponding to the query Q, key K, and value V of the i-th head; this matrix is continuously updated during training; In the multi-head attention mechanism, the calculation expression of each head's attention is as follows: head i = Attention(Q i , K i , V i ) (18) After calculating all single-head attentions, the results of each head's attention are concatenated, and then multiplied by the weights of each head to output the result of the multi-head attention mechanism. The expression is as follows: MHA(Q, K, V) = Concat(head1, head2, …, head N )W O (19) Among them, W O represents the output weight of the multi-head self-attention, and head i represents the i-th head; S34. Based on steps S31 - S33, use the ADM-MATD3 reinforcement learning algorithm to solve the strategy of multi-agent cooperative deception jamming networking radar; S341. Initialize the coordinates of the multi-agent cluster and the coordinates of the networking radar, set the number of randomly sampled experiences for each single reinforcement learning training, initialize the number of delayed update rounds and the maximum number of training rounds; S342. Initialize the value network, policy network parameters of all agents and the experience replay pool, and reset the environment to obtain the initial state s0; S343. Each agent independently observes the environment to obtain the current state After each agent inputs its own state into its respective policy network, it outputs an action a for each agent, and a superimposes perturbations according to the dynamic search rate: Among them, represents the output obtained after the current state is input into the policy network, ε represents the Gaussian noise that conforms to the dynamic exploration variance of the current round, and std(episode) represents the dynamic exploration variance of the current round calculated in formula (14); S345. The agent executes action a, and the environment performs a state transition according to the agent's action, outputs a reward and the state s' at the next moment, and stores the obtained experience in the experience pool; S346. If the delayed update condition is met, enter step S347; otherwise, jump back to step S343; Among them, the delayed update means that after every preset number of update rounds, the agent conducts neural network training and gradient descent update, and during the training rounds, the agent only conducts exploration and experience accumulation; S347. Based on the preset number of randomly sampled experience values in step S341, take a fixed amount of experience from the experience pool to train the value network and the policy network; S348. Use the gradient descent method to update the policy network and the soft update method to update the target network; S349. If the number of training rounds is equal to the set maximum number of training rounds, save the ADM-MATD3 network model; otherwise, reset the environment, and loop through steps S344 - S349 until the set number of training rounds is reached.

Citation Information

Cited By

  • Cognitive interference decision-making method based on deep reinforcement learning and application system thereof

    CN121142484A