A multi-agent reinforcement learning-based multi-domain joint interference resource allocation method and system

By constructing a multi-agent reinforcement learning-based model for a multi-to-multi adversarial environment and designing a global reward function, this approach addresses the issues of decision delay and insufficient resource management in the allocation of jamming resources in complex battlefield environments. It achieves efficient and flexible allocation of jamming resources, thereby improving the jamming efficiency and adaptability against enemy radar systems.

CN119789144BActive Publication Date: 2025-10-21HARBIN ENG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411961184.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-10-21
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

Existing jamming resource allocation technologies have limitations in the face of the complexity and dynamism of modern electronic warfare. They are difficult to achieve efficient and flexible resource management and allocation in environments with multiple targets and multiple jamming sources. Traditional methods exhibit decision delays, slow response speeds, resource constraints, and insufficient energy management in complex and ever-changing battlefield environments.

Method used

A multi-agent reinforcement learning-based approach is adopted to construct a multi-to-multi adversarial environment model, define the joint state space and action space of multiple jammers, design a global reward function, and use a value decomposition network algorithm to learn the optimal policy, thereby realizing the dynamic adjustment of jamming beams and power, and resource allocation through joint representation of multi-domain information.

Benefits of technology

It improves the efficiency and flexibility of jamming enemy radar systems, has a high real-time response capability, can make rapid decisions in complex and ever-changing battlefield environments, enhances the robustness and adaptability of the model, and simplifies the model implementation and maintenance process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119789144B_ABST
    Figure CN119789144B_ABST
Patent Text Reader

Abstract

The application discloses a multi-agent reinforcement learning-based multi-domain joint interference resource allocation method and system, wherein the method steps comprise the following steps: based on a multi-jammer cooperative jamming task, a multi-to-multi confrontation environment model is constructed; based on the multi-to-multi confrontation environment model, a multi-jammer joint state space is defined; based on the multi-to-multi confrontation environment model, a multi-jammer joint action space is designed; based on the multi-jammer joint state space and the multi-jammer joint action space, a global reward function of multi-domain information joint representation is constructed; based on the global reward function, optimal policy learning is carried out; and a multi-agent system makes a decision according to the learned optimal policy. Through the value decomposition network algorithm, the multi-jammer joint state space, the multi-jammer joint action space and the global reward function are designed, the dynamic adjustment of jamming beam allocation and jamming power size of the multi-jammer of our side is realized, and therefore the jamming efficiency and flexibility on the enemy radar system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic countermeasures, and in particular to a multi-domain joint interference resource allocation method and system based on multi-agent reinforcement learning. Background Art

[0002] With the development of modern electronic warfare, confrontation scenarios have become increasingly complex, evolving from early "one-on-one" confrontations to complex "many-on-many" situations. With the rapid increase and diversification of electronic equipment in battlefield environments, enemy and friendly electronic warfare systems must not only confront multiple targets but also respond rapidly in a highly dynamic and uncertain environment. This complex confrontation scenario with multiple targets and interference sources poses unprecedented challenges to the management and allocation of interference resources.

[0003] The jamming resource allocation problem focuses on how to optimally utilize limited resources, such as the number of jammers, jamming power, and jamming frequency range, to achieve maximum jamming effects on enemy electronic systems. Traditionally, this process often relies on fixed preset rules, past experience, or optimization algorithms to make decisions. However, as the electronic warfare environment becomes more complex and rapidly changing, these methods are gradually showing their lack of flexibility and adaptability.

[0004] Reinforcement learning (RL) is a machine learning method that learns optimal strategies through the interaction between an intelligent agent and its environment. The characteristic of RL is its ability to autonomously learn optimal behavioral strategies through trial and error and reward mechanisms without explicit guidance. In recent years, RL has achieved remarkable success in many fields, especially in dealing with complex, dynamic, and uncertain environments. However, traditional single-agent RL algorithms have obvious shortcomings when dealing with "many-to-many" adversarial scenarios. First, single-agent systems need to centrally process large amounts of information, resulting in decision delays and slow response speeds. Second, single-agent systems have difficulty adapting to complex and changing network topologies and cannot effectively cope with the needs of multi-objective optimization and adversarial interference. In addition, single-agent systems are also unable to cope with resource constraints and energy consumption management, making it difficult to maintain efficient operation for a long time in high-intensity adversarial environments.

[0005] Multi-Agent Reinforcement Learning (MARL) further expands the application scope of reinforcement learning, enabling it to address collaboration between multiple agents. Through distributed decision-making and parallel processing, multi-agent systems can rapidly adapt to complex and changing battlefield environments and achieve efficient and flexible resource allocation. Each agent can make independent decisions based on its own state and that of its surroundings, while simultaneously optimizing the overall jamming effect through collaboration and competition. Furthermore, multi-agent systems possess high robustness and fault tolerance, enabling them to maintain stable operation even if some agents fail.

[0006] Numerous researchers have studied the multi-jammer interference resource allocation problem. Shen Yang et al. transformed the radar jamming resource optimization problem into a 0-1 programming problem and applied the Hungarian method to solve the problem and obtain an interference resource allocation strategy. Zhang et al. proposed a two-step solution based on particle swarm optimization to solve the joint interference beam and power allocation problem. You et al. established a threat assessment and interference allocation model based on combinatorial optimization and proposed a differential evolution algorithm based on extended permutation to optimize the interference coding matrix, effectively reducing the threat posed by networked radars to targets under multiple constraints. Jiang et al. proposed a hybrid quantum-behavior particle swarm optimization and self-tuning genetic algorithm to optimize interference resource allocation under multiple constraints. Yao et al. proposed an improved firefly algorithm to optimize the interference resource allocation model and used random keys to improve the encoding method of the firefly algorithm. Qi et al. constructed an interference resource allocation model that considered information in the spatial, frequency, and energy domains. They used the Determined Boolean Order (DBO) algorithm and the Q-learning algorithm to optimize the interference beam allocation problem and the interference power allocation problem, respectively, achieving good convergence and timeliness. Panzes et al. used the interference-to-signal ratio as an indicator for evaluating jamming effectiveness and employed a multi-agent reinforcement learning deep neural network algorithm consisting of a performer network and a critic network to achieve collaborative jamming resource allocation for multiple jammers. Existing research on jamming resource allocation has mostly focused on a single domain or a combination of these domains: spatial, frequency, and energy. However, multi-domain information has not been jointly processed within a unified framework. This means that existing solutions may have limitations when dealing with complex multi-domain environments, as they often view the role of each domain in isolation rather than optimizing resource allocation holistically. Furthermore, as the scale of jamming resource allocation expands, traditional algorithms and single-agent reinforcement learning algorithms are prone to problems such as dimensionality explosion, slower convergence, reduced optimization probability, and longer algorithm response times, making it difficult to obtain optimal solutions and meeting practical application needs.

[0007] In summary, existing jamming resource allocation technologies have many limitations when facing the complexity and dynamics of modern electronic warfare. Therefore, a new approach is urgently needed to address these issues. Summary of the Invention

[0008] To solve the technical problems in the above background, the present invention aims to improve the efficiency and flexibility of jamming enemy radar systems by dynamically adjusting the jamming beam allocation and power of our jammers; at the same time, it realizes efficient jamming resource allocation in an intelligent manner to enhance the overall effectiveness of electronic warfare.

[0009] To achieve the above objectives, the present invention provides a multi-domain joint interference resource allocation method based on multi-agent reinforcement learning, comprising the following steps:

[0010] Based on the multi-jammer coordinated jamming task, a many-to-many confrontation environment model is constructed;

[0011] Based on the many-to-many confrontation environment model, defining a multi-jammer joint state space;

[0012] Based on the many-to-many confrontation environment model, a multi-jammer joint action space is designed;

[0013] Based on the multi-disruptor joint state space and the multi-disruptor joint action space, constructing a global reward function for joint representation of multi-domain information;

[0014] Based on the global reward function, optimal strategy learning is performed;

[0015] The multi-agent system makes decisions based on the learned optimal strategy.

[0016] Preferably, the method for constructing the many-to-many confrontation environment model includes: defining each of our jammers as an independent intelligent agent, specifically capable of perceiving the state of the environment, executing actions, and receiving rewards; having a communication and collaboration mechanism between the jammers for sharing information or coordinating action strategies; and the confrontation environment contains the position information X of each jammer. j , the maximum output power P of each jammer j_max , the maximum number of assignable beams N beam , interference frequency range W j , airspace condition constraints K, location information of each radar X r , the operating frequency range of each radar W r ; Construct interference parameter matrix M j =[X j ,W j ,P j_max ] and radar parameter matrix M r =[X r ,W r] to represent environmental information.

[0017] Preferably, the method for defining the multi-jammer joint state space includes: encoding state elements in matrix form so that the multi-jammers can better understand and process complex state information; the state elements include: the interference beam distribution relationship and the interference power distribution relationship between all jammers and radars in the current confrontation round n.

[0018] Preferably, multiple jammers work together n The method for designing the joint action space of multiple jammers includes: discretizing the continuous action space of each jammer into l action options, so as to reduce the complexity of the action space of each jammer while maintaining sufficient flexibility; in this way, multiple jammers are helped to select the optimal action combination among the limited action options.

[0019] Preferably, the method based on the multi-jammer joint state space and the multi-jammer joint action space includes: combining spatial domain, frequency domain and energy domain information to design a unified global reward function to guide the jammer to learn the optimal strategy; the global reward function integrates multiple evaluation factors, including the frequency domain overlap between the jammer and the radar, the distance between the jammer and the radar, the degree of influence on the signal-to-interference ratio of the radar receiver, and other constraints.

[0020] Preferably, the VDN algorithm is used to perform the optimal strategy learning, and the process includes:

[0021] The value decomposition network algorithm is used to train the multi-agent system. The network structure of each jammer consists of two neural networks: an estimated Q network for estimating the current strategy and a target Q network for stabilizing training. The jammer performs actions according to the current strategy and collects four-tuples (s n ,a n ,R n ,s n+1 ), stored in the experience replay buffer; when the buffer is full, sample some data from it, calculate the TD error, and update the parameters of the estimated Q network through the back propagation algorithm; regularly copy the parameters of the estimated Q network to the target Q network to maintain the stability of the learning process; adopt the ε-greedy strategy to gradually reduce the proportion of random behavior.

[0022] The present invention also provides a multi-domain joint interference resource allocation system based on multi-agent reinforcement learning, which is used to implement the above method and includes: a building module, a definition module, a design module, a construction module, a learning module, and a decision module;

[0023] The building module is used to build a many-to-many confrontation environment model based on a multi-jammer collaborative jamming task;

[0024] The definition module is used to define a multi-jammer joint state space based on the many-to-many confrontation environment model;

[0025] The design module is used to design a multi-jammer joint action space based on the many-to-many confrontation environment model;

[0026] The construction module is used to construct a global reward function of joint representation of multi-domain information based on the multi-disruptor joint state space and the multi-disruptor joint action space;

[0027] The learning module is used to perform optimal strategy learning based on the global reward function;

[0028] The decision-making module is used by the multi-agent system to make decisions based on the learned optimal strategy.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] (1) By adopting the value decomposition network (VDN) algorithm and designing the joint state space, joint action space, and global reward function of multiple jammers, the dynamic adjustment of the jamming beam allocation and jamming power of our multiple jammers is achieved, thereby improving the jamming efficiency and flexibility of the enemy radar system and ensuring efficient jamming in a complex and changing battlefield environment. In addition, this method has a high degree of real-time response capability and can make decisions quickly in a short period of time to adapt to the rapid changes in the battlefield environment. Through adaptive learning, the model can continuously enhance its robustness and adaptability, reduce its dependence on external conditions, and simplify the model implementation and maintenance process.

[0031] (2) By constructing a global reward function that jointly represents multi-domain information, and taking into account multiple factors such as the frequency domain overlap between the jammer and the radar, the distance between the jammer and the radar, the degree of impact on the signal-to-interference ratio of the radar receiver, and other constraints, we can achieve an overall optimized allocation of our limited jamming resources from multiple dimensions, enhance the reliability and operability of the model, and provide strong technical support for modern electronic warfare. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention;

[0034] Figure 2 Schematic diagram of the optimal strategy learning process according to an embodiment of the present invention;

[0035] Figure 3 Detailed flowchart of an embodiment of the present invention;

[0036] Figure 4 This is a schematic diagram of visualizing spatial situation information before interference resource allocation according to an embodiment of the present invention;

[0037] Figure 5 This is a diagram showing the change in reward value for a single round of the VDN algorithm according to an embodiment of the present invention;

[0038] Figure 6 This is a visualization diagram of the result after interference resource allocation according to an embodiment of the present invention. DETAILED DESCRIPTION

[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0040] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] Example 1

[0042] like Figure 1 FIG. 1 is a flow chart of the method of this embodiment, and the steps include:

[0043] S1. Based on the multi-jammer collaborative jamming task, a many-to-many confrontation environment model is constructed.

[0044] Multiple multi-beam jammers in space perform coordinated jamming tasks against multiple radars on the ground. Each jammer's multi-beam jamming system can simultaneously generate multiple beams to interfere with multiple radars in different directions, and resources such as the jamming beam's direction, number, and transmit power can be flexibly controlled. Because each jammer can provide limited jamming power, to improve the utilization of friendly jamming resources, it is necessary to rationally allocate the beam direction of each jammer and the transmit power of different beams based on real-time battlefield situation information. This way, limited jamming resources can be used to achieve the highest possible jamming benefit. Therefore, this embodiment constructs a "many-to-many" confrontation environment model based on real-world space scenarios.

[0045] Specifically, each jammer on our side is defined as an independent intelligent agent with the ability to perceive the environment, perform actions, and receive rewards. A certain communication and collaboration mechanism is required between jammers to share information or coordinate action strategies. The adversarial environment contains the location information of each jammer X j , the maximum output power P of each jammer j_max , the maximum number of assignable beams N beam , interference frequency range W j , airspace condition constraints K, location information of each radar X r , the operating frequency range of each radar W r . Construct the interference parameter matrix M j =[X j ,W j ,P j_max ] and radar parameter matrix M r =[X r ,W r ] to represent environmental information.

[0046] S2. Based on the many-to-many confrontation environment model, define the joint state space of multiple jammers.

[0047] In the constructed many-to-many confrontation environment model, the joint state of multiple jammers s n The system consists of two state elements: the jamming beam distribution relationship between all jammers and radars in the current confrontation round n, and the jamming power distribution relationship. These state elements are encoded in a matrix form. This encoding method helps multiple jammers better understand and process complex state information.

[0048] Specifically, assume that there are I radars and J multi-beam jammers in the current confrontation environment. Each jammer can simultaneously allocate multiple jamming beams to interfere with multiple radars. In the current confrontation round n, the jamming beam pointing distribution relationship of all jammers to radars can be expressed by a binary variable matrix U n To characterize:

[0049]

[0050] in, Represents a binary variable that can only take the value "0" or "1"; Indicates that in the nth confrontation round, jammer j allocates beams to jam radar i; Indicates that jammer j does not allocate beam to jam radar i.

[0051] For the transmission power allocation of different interference beams, it can be represented by a matrix P n To characterize:

[0052]

[0053] in, It represents the transmission power of the jammer beam assigned to radar i by jammer j in the nth confrontation round.

[0054] Therefore, the local observation state of each jammer j is Multi-jammer joint state The multi-jammer joint state space S of the present invention can be expressed as:

[0055] S={s1,s2,…,s n ,…,s N} (3)

[0056] Among them, N is the total number of confrontation rounds.

[0057] S3. Based on the many-to-many adversarial environment model, design the joint action space of multiple jammers.

[0058] In the constructed many-to-many confrontation environment model, multiple jammers jointly act a n is the combination of actions taken by all jammers in the current round n of the game. Discretizing each jammer's continuous action space into l action options reduces the complexity of each jammer's action space while maintaining sufficient flexibility. In this way, multiple jammers can select the optimal action combination from a limited set of action options.

[0059] Specifically, in the nth round of confrontation, the action of jammer j is in, The initial action is randomly selected from these l discrete values, and the next action is selected according to the model's strategy π.

[0060] Therefore, multiple jammers work together The multi-jammer joint action space A of the present invention can be expressed as:

[0061] A={a1,a2,...,a n ,…,a N} (4)

[0062] S4. Based on the joint state space and joint action space of multiple distractors, a global reward function that jointly represents multi-domain information is constructed.

[0063] A unified global reward function, combining spatial, frequency, and energy domain information, is designed to guide the jammer in learning the optimal strategy. To this end, the global reward function incorporates multiple evaluation factors, including the frequency domain overlap between the jammer and the radar, the distance between the jammer and the radar, the impact on the radar's signal-to-interference ratio (SIR) at the receiver, and other constraints. This global reward function allows the jammer to adjust its behavior based on different mission requirements, thereby achieving multi-objective optimization.

[0064] (1) Constraints

[0065] ① Airspace condition constraints. Define a binary variable matrix K to characterize the airspace position relationship between the jammer and the radar:

[0066]

[0067] Among them, k ji =1 means that jammer j is within the beam coverage of radar i and can be assigned to jam the radar i. ji =0 means that jammer j is not within the beam coverage of radar i and cannot be assigned to jam this radar.

[0068] ② Interference beam allocation constraint. Assume that each jammer can allocate at most h beams to interfere with multiple radars at the same time, and each radar can only be allocated at most one interference beam:

[0069]

[0070] ③ Interference power allocation constraints. Each radar being interfered with must be allocated interference power. It is assumed that when the SIR at the radar end is greater than 14dB, it means that the interference of the own side is completely ineffective against the radar. When the SIR is less than 3dB, it means that the radar has been completely and effectively interfered. Therefore, there is no need to continue to allocate interference power outside this range, otherwise it will cause a waste of interference resources. At the same time, the sum of the transmission power allocated to each interference beam of the multi-beam jammer cannot exceed the maximum interference power that the jammer can provide:

[0071]

[0072] Define a reward term V that represents whether each constraint condition is satisfied. The expression is:

[0073]

[0074] (2) Objective function

[0075] ① Frequency domain overlap between the jammer and the radar. Assume that the jamming frequency range of the jammer j is [f j1 ,f j2 ], the operating frequency range of radar i is [f i1 ,f i2], then the frequency domain overlap between jammer j and radar i is r ji The calculation formula is:

[0076]

[0077] Among them, B i represents the operating bandwidth of radar i, Δf ji =min(f j2 ,f i2 )-max(f j1 ,f i1 ),

[0078] Therefore, the frequency domain overlap between the jammer and the radar in the entire jamming resource allocation scheme is C f :

[0079]

[0080] Where α1 represents the normalization coefficient.

[0081] ② The distance between the jammer and the radar. Assume that the position of the jammer j is (x j ,y j ,z j ), the position of radar i is (x i ,y i ,z i ), then the distance d between the two ji It can be expressed as:

[0082]

[0083] Therefore, the total distance D between the jammer and the radar in the entire jamming resource allocation scheme is:

[0084]

[0085] Wherein, α2 represents the normalization coefficient.

[0086] ③ The impact on the signal-to-interference ratio (SIR) of the radar receiver. The purpose of jammer resource allocation is to select appropriate jamming power for the jammer to reduce the SIR of the radar receiver. However, in reality, the jammer cannot directly obtain the SIR of the radar receiver. Therefore, the SIR of the radar receiver can be estimated based on the SIR of the jammer.

[0087] In the confrontation scenario model established by the present invention, the jammer's purpose of jamming the radar is to shield itself from being detected by the radar. Therefore, the distance between the radar and the target is equal to the distance between the radar and the jammer. Based on the radar equation, the signal-to-interference ratio at the radar receiver can be calculated as:

[0088]

[0089] Where σ represents the target radar cross-sectional area; P i Indicates the transmission power of the radar signal; d ji Indicates the distance between the radar and the jammer; G i and G j They represent the power gains of the radar and jammer antennas respectively; μ represents the polarization matching loss coefficient between the jammer signal and the radar signal.

[0090] Similarly, the signal-to-interference ratio at the jammer end can be calculated as:

[0091]

[0092] Where λ is the wavelength of the transmitted signal.

[0093] Then the ratio of the signal-to-interference ratio at the radar receiver to the signal-to-interference ratio at the jammer is:

[0094]

[0095] Therefore, the signal-to-interference ratio SIR at the jammer end can be jammer Estimate the signal-to-interference ratio (SIR) at the radar receiver radar :

[0096]

[0097] In order to evaluate the transmit power of jammer j, The influence of the interference signal on the signal-to-interference ratio of the radar i receiving end is defined as a variable η ji To measure:

[0098]

[0099] Where δ represents the proportional factor, which is related to the minimum signal-to-interference ratio that the jammer expects to achieve at the radar receiver.

[0100] Therefore, the impact of the entire interference resource allocation scheme on the signal-to-interference ratio of the radar receiver is:

[0101]

[0102] Wherein, α3 represents the normalization coefficient.

[0103] In summary, the global reward function formula is:

[0104] R=(ω1·C f +ω2·D+ω3·J p +ρ·V)·β (19)

[0105] Among them, ω1, ω2, and ω3 represent weight coefficients, which are used to balance the importance of different goals; ρ represents the reward factor that satisfies the constraints, and ω1+ω2+ω3+ρ=1; β is the reward function magnification factor.

[0106] S5. Based on the global reward function, learn the optimal strategy.

[0107] This embodiment implements the VDN algorithm to learn the optimal strategy. The Value Decomposition Network (VDN) algorithm is used to train the multi-agent system. The network structure of each jammer consists of two neural networks: one is the estimated Q network for estimating the current strategy, and the other is the target Q network for stabilizing training. The jammer performs actions according to the current strategy and collects four-tuples (s n ,a n ,R n ,s n+1 ) and stored in the experience replay buffer. When the buffer is full, a batch of data is sampled from it, the TD error is calculated, and the parameters of the estimated Q network are updated using the backpropagation algorithm. The parameters of the estimated Q network are periodically copied to the target Q network to maintain the stability of the learning process. An ε-greedy strategy or other methods are used to balance exploration and exploitation, gradually reducing the proportion of random behavior. Through this approach, the multi-agent system can learn the optimal interference resource allocation strategy.

[0108] Specifically, the process of implementing the VDN algorithm to learn the optimal strategy is as follows: Figure 2 In the VDN algorithm, the network structure of each jammer contains two neural networks: one is the estimated Q network used to estimate the current strategy, and the other is the target Q network used for stable training.

[0109] (1) Value decomposition

[0110] VDN assumes that the joint Q-value function can be obtained by the local Q-value function of each jammer Perform linear decomposition:

[0111]

[0112] in, represents the local observation of jammer j; represents the action of jammer j; Indicates that jammer j is based on its own local observation and actions The learned local Q value; Q tot represents the global reward function.

[0113] This linear decomposition allows each jammer to independently choose actions during execution, while optimizing the strategy through a global Q-value function during centralized training.

[0114] (2) Training process

[0115] VDN training adopts a centralized training and decentralized execution mode.

[0116] Centralized training: During training, you can access the global information of all jammers, such as the global state s n and joint action a n , using this information to calculate the global reward function. At the same time, the joint Q-value function is calculated and updated by the sum of the local Q-value functions.

[0117] Decentralized execution: During the execution phase, the jammer can only rely on its own local information And the learned local Q value function Make action selection.

[0118] In this way, each jammer can perform independently while ensuring the learning of the global optimal solution during the training phase.

[0119] (3) Loss function

[0120] During training, the loss function of VDN is similar to traditional Q-learning, which updates the Q value based on the TD error. For a batch of data sampled from the experience replay buffer, the loss function is:

[0121]

[0122] Target Q value y b The calculation is as follows:

[0123]

[0124] Where B represents the batch size; (s b ,a b ,R b ,s b+1 ) represents the sample sampled from the experience replay buffer; R b Represents the global reward value given by the environment; γ represents the discount factor; s b+1 Indicates the next state; a b+1 represents the optimal joint action in the next state.

[0125] Because Q tot (s n ,a n ) is calculated by summing each local Q value, and updating Q tot At the same time, the local Q value Q of each jammer will be updatedj .

[0126] (4) Exploration and Utilization:

[0127] Use ε-greedy strategy or other methods to balance exploration and exploitation, and gradually reduce the proportion of random behavior:

[0128]

[0129] Among them, ε represents the exploration probability.

[0130] S6. The multi-agent system makes decisions based on the learned optimal strategy.

[0131] When the Q network converges, the jammer selects the optimal action based on the learned optimal strategy. The jammer continuously senses environmental changes and responds quickly and optimally based on the latest state. At a certain moment, the multi-agent system selects the optimal action combination a* based on the current state s and the learned optimal strategy:

[0132]

[0133] In actual operation, the jammer needs to constantly sense environmental changes and respond quickly based on the latest status. In this way, the multi-agent system can make autonomous decisions in complex battlefield environments, ultimately achieving efficient, flexible, and reliable jamming resource allocation. The detailed process of this embodiment is as follows: Figure 3 shown.

[0134] Example 2

[0135] The accuracy of the present invention will be described in detail below in conjunction with the effectiveness verification of this embodiment.

[0136] (1) Feasibility verification

[0137] During the feasibility verification phase, a complex simulation confrontation scenario of “multiple jammers against multiple radars” was set up. It was assumed that at a certain moment in the confrontation space, there were J = 3 friendly jammers performing coordinated jamming tasks against I = 5 enemy radars. The radar parameter information and jammer parameter information at this time are shown in Table 1 and Table 2, respectively.

[0138] Table 1

[0139]

[0140] Table 2

[0141]

[0142] Visualization of spatial situation information before interference resource allocation Figure 4 shown.

[0143] The spatial condition constraint matrix at this time is shown in Table 3.

[0144] Table 3

[0145]

[0146] Other parameter settings for the adversarial scenario are shown in Table 4.

[0147] Table 4

[0148]

[0149] The change of the single-round reward value of the VDN algorithm in this embodiment is as follows: Figure 5 As shown in the figure, as the model and the environment continue to interact, the jammer continuously learns the experience of allocating jamming resources that can maximize the reward value, and the algorithm reaches convergence around 120 rounds. In the vast majority of subsequent rounds, the jammer can select the action that can maximize the global reward value based on the learned strategy. In a few rounds, the jammer randomly selects actions. This is due to the use of the ε-greedy strategy, which allows the jammer to maintain a certain exploration rate in the later stages of training to avoid falling into local optimality.

[0150] The optimal multi-jammer cooperative interference resource allocation scheme obtained in this embodiment is shown in Table 5 and Table 6. The visualization results are shown in Table 5 and Table 6. Figure 6 As shown in the figure. According to this scheme, all five radars in the countermeasure scenario were assigned jamming beams for interference. Taking the first jammer as an example, it was assigned to interfere with the third and fifth radars, with the transmission powers of the two jamming beams being 0.12kW and 0.08kW, respectively. The experimental results show that the multi-domain joint jamming resource allocation model based on multi-agent reinforcement learning proposed in this paper has certain feasibility.

[0151] Table 5

[0152]

[0153] Table 6

[0154]

[0155] (2) Method comparison and performance analysis

[0156] In the same simulation scenario, the VDN algorithm used in this embodiment was compared with the Dung Beetle Optimization (DBO) algorithm and the Particle Swarm Optimization (PSO) algorithm in terms of interference benefit and algorithm response time. To ensure the accuracy and fairness of the comparison, the operating parameters of the jammer and radar were kept consistent. Each algorithm was subjected to 200 Monte Carlo experiments under the same hardware conditions: a 13th-generation i5-13600KF CPU, a 3060Ti GPU, 32GB of RAM, and Windows 11. The average normalized interference benefit and algorithm response time of the different algorithms are shown in Table 7.

[0157] Table 7

[0158]

[0159] Compared with the DBO algorithm and the PSO algorithm, in terms of interference benefit, the interference benefit obtained by the VDN algorithm increased by 13.95% and 16.67% respectively; in terms of algorithm response time, the response time of the VDN algorithm was shortened by 98.31% and 97.99% respectively, indicating that the model has good effectiveness and timeliness.

[0160] The method proposed in this embodiment uses the Value Decomposition Network (VDN) algorithm to design a joint state space and action space for multiple jammers, and constructs a global reward function that jointly represents multi-domain information. This method comprehensively considers multiple factors, including the frequency domain overlap between the jammer and the radar, the distance between the jammer and the radar, the impact on the radar receiver's signal-to-interference ratio, and other constraints, thereby achieving a comprehensive and multi-dimensional optimization of our limited jamming resources. Experimental results show that compared to the other two swarm intelligence optimization algorithms, the method proposed in this embodiment can achieve higher jamming benefits and has a significant advantage in algorithm response time.

[0161] Example 3

[0162] This embodiment also provides a multi-domain joint interference resource allocation system based on multi-agent reinforcement learning, including: a building module, a definition module, a design module, a construction module, a learning module, and a decision module; the building module is used to construct a multi-to-many adversarial environment model based on the multi-jammer collaborative interference task; the definition module is used to define the multi-jammer joint state space based on the multi-to-many adversarial environment model; the design module is used to design the multi-jammer joint action space based on the multi-to-many adversarial environment model; the construction module is used to construct a global reward function for the joint representation of multi-domain information based on the multi-jammer joint state space and the multi-jammer joint action space; the learning module is used to perform optimal strategy learning based on the global reward function; and the decision module is used for the multi-agent system to make decisions based on the learned optimal strategy.

[0163] The workflow of the building module includes: defining each jammer as an independent intelligent agent with the ability to perceive the environment state, perform actions, and receive rewards; jammers have a communication and collaboration mechanism to share information or coordinate action strategies; the adversarial environment contains the location information of each jammer X j , the maximum output power P of each jammer j_max , the maximum number of assignable beams N beam , interference frequency range W j , airspace condition constraints K, location information of each radar X r , the operating frequency range of each radar W r ; Construct interference parameter matrix M j =[X j ,W j ,P j_max ] and radar parameter matrix M r =[X r ,W r ] to represent environmental information.

[0164] The workflow of the definition module includes: encoding state elements in matrix form for multiple jammers to better understand and process complex state information; state elements include: the interference beam distribution relationship and interference power distribution relationship of all jammers and radars in the current confrontation round n.

[0165] The workflow of the design module includes: multi-jammer joint action a n It is formed by the combination of actions taken by all jammers in the current confrontation round n; the method for designing the joint action space of multiple jammers includes: discretizing the continuous action space of each jammer into l action options, which is used to reduce the complexity of the action space of each jammer while maintaining sufficient flexibility; in this way, multiple jammers are helped to select the optimal action combination among the limited action options.

[0166] The workflow of the construction module includes: combining spatial, frequency, and energy domain information to design a unified global reward function to guide the jammer to learn the optimal strategy; the global reward function integrates multiple evaluation factors, including the frequency domain overlap between the jammer and the radar, the distance between the jammer and the radar, the degree of impact on the signal-to-interference ratio of the radar receiver, and other constraints.

[0167] The workflow of the learning module includes: using the value decomposition network algorithm to train the multi-agent system, and the network structure of each jammer contains two neural networks: an estimated Q network for estimating the current strategy, and a target Q network for stabilizing training; the jammer performs actions according to the current strategy and collects four-tuples (s n ,a n ,R n ,sn+1 ), stored in the experience replay buffer; when the buffer is full, sample some data from it, calculate the TD error, and update the parameters of the estimated Q network through the back propagation algorithm; regularly copy the parameters of the estimated Q network to the target Q network to maintain the stability of the learning process; adopt the ε-greedy strategy to gradually reduce the proportion of random behavior.

[0168] The decision-making module's workflow involves the following: After the Q network converges, the jammer selects the optimal action based on the learned optimal strategy. The jammer continuously senses environmental changes and rapidly responds to the latest state. At a specific moment, the multi-agent system selects the optimal action combination based on the current state and the learned optimal strategy. In actual operation, the jammer needs to continuously sense environmental changes and respond quickly based on the latest state. This approach enables the multi-agent system to make autonomous decisions in complex battlefield environments, ultimately achieving efficient, flexible, and reliable jamming resource allocation.

[0169] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A multi-domain joint interference resource allocation method based on multi-agent reinforcement learning, characterized in that the steps include: Based on the multi-jammer coordinated jamming task, a many-to-many confrontation environment model is constructed; Based on the many-to-many confrontation environment model, defining a multi-jammer joint state space; The state elements include: the interference beam distribution relationship and the interference power distribution relationship of all jammers and radars in the current confrontation round n; Based on the many-to-many confrontation environment model, a multi-jammer joint action space is designed; a multi-jammer joint action space is designed. n It is formed by the combination of actions taken by all jammers in the current confrontation round n; Based on the multi-jammer joint state space and the multi-jammer joint action space, a global reward function for joint representation of multi-domain information is constructed; the global reward function formula is: R=(ω1·C f +ω2·D+ω3·J p +ρ·V)·β Among them, ω1, ω2, and ω3 represent weight coefficients, which are used to balance the importance of different goals; ρ represents the reward factor that satisfies the constraints, and ω1+ω2+ω3+ρ=1; β is the reward function magnification; C f represents the frequency domain overlap between the jammer and the radar in the entire jamming resource allocation scheme; D represents the total distance between the jammer and the radar in the entire jamming resource allocation scheme; J p It represents the influence of the entire interference resource allocation scheme on the signal-to-interference ratio of the radar receiver; V represents a reward item that characterizes whether each constraint condition is met, and its expression is: The constraints include: Airspace condition constraint, by defining a binary variable matrix K to characterize the airspace position relationship between the jammer and the radar; k ji =1 means that jammer j is within the beam coverage of radar i and can be assigned to jam the radar i. ji =0 means that jammer j is not within the beam coverage of radar i and cannot be assigned to jam this radar; Jamming beam allocation constraint: it is assumed that each jammer can allocate at most h beams to jam multiple radars simultaneously, and each radar can only be allocated at most one jamming beam; Interference power allocation constraints, the expressions include: Among them, P ji Indicates the SIR of the radar end; Indicates the transmit power of the jammer; P j_max Indicates the maximum output power of each jammer; represents a binary variable that can only take the value "0" or "1"; J represents the total number of multi-beam jammers; Based on the global reward function, optimal strategy learning is performed; The multi-agent system makes decisions based on the learned optimal strategy.

2. The multi-domain joint interference resource allocation method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The method for constructing the multi-to-multi adversarial environment model includes: defining each jammer on our side as an independent intelligent agent with the ability to perceive the environment state, perform actions, and receive rewards; the jammers have a communication and collaboration mechanism for sharing information or coordinating action strategies; the adversarial environment contains the position information X of each jammer j , the maximum output power P of each jammer j_max , the maximum number of assignable beams N beam , interference frequency range W j , airspace condition constraints K, location information of each radar X r , the operating frequency range of each radar W r ; Construct interference parameter matrix M j =[X j ,W j ,P j_max ] and radar parameter matrix M r =[X r ,W r ] to represent environmental information.

3. The multi-domain joint interference resource allocation method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The method for defining the multi-jammer joint state space includes: encoding state elements in a matrix form, so that the multi-jammers can better understand and process complex state information.

4. The multi-domain joint interference resource allocation method based on multi-agent reinforcement learning according to claim 3 is characterized in that: The method for designing the joint action space of multiple jammers includes: discretizing the continuous action space of each jammer into l action options, so as to reduce the complexity of the action space of each jammer while maintaining sufficient flexibility; in this way, multiple jammers are helped to select the optimal action combination among limited action options.

5. The multi-domain joint interference resource allocation method based on multi-agent reinforcement learning according to claim 4 is characterized in that: The method based on the multi-jammer joint state space and the multi-jammer joint action space includes: combining spatial domain, frequency domain and energy domain information to design a unified global reward function for guiding the jammer to learn the optimal strategy.

6. The multi-domain joint interference resource allocation method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The optimal strategy learning is performed using a value decomposition network algorithm, and the process includes: The value decomposition network algorithm is used to train the multi-agent system. The network structure of each jammer consists of two neural networks: an estimated Q network for estimating the current strategy and a target Q network for stabilizing training. The jammer performs actions according to the current strategy and collects four-tuples (s n ,a n ,R n ,s n+1 ), where s n is the joint state of multiple jammers, a n The proposed method is based on the joint action of multiple jammers, R is the global reward function, and is stored in the experience replay buffer. When the buffer is full, some data is sampled from it, the temporal difference error is calculated, and the parameters of the estimated Q network are updated through the back propagation algorithm. The parameters of the estimated Q network are periodically copied to the target Q network to maintain the stability of the learning process. An ε-greedy strategy is adopted to gradually reduce the proportion of random behavior.

7. A multi-domain joint interference resource allocation system based on multi-agent reinforcement learning, characterized in that: include: Building module, defining module, design module, construction module, learning module, decision module, the system is used to implement the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Radar anti-interference intelligent decision-making method based on reinforcement learning

    CN113625233A

  • Rader system

    JP2013130527A