A Networked Radar Decision-Making Method Based on Improved MADDPG, Balancing Resolution and Interference Resistance
By constructing a continuous transmission and reception signal model for networked radars, and combining a multi-agent Markov decision process and a deep deterministic policy gradient algorithm, the problems of broadband frequency sweeping interference and adjacent frequency interference in multi-radar networking scenarios are solved, thereby improving high-resolution detection and anti-interference capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHEASTERN UNIV AT QINHUANGDAO
- Filing Date
- 2026-04-23
- Publication Date
- 2026-05-26
AI Technical Summary
In multi-radar networking scenarios, there is still a lack of effective solutions for simultaneously addressing broadband frequency sweep interference avoidance and adjacent frequency interference reduction through collaborative decision-making in the continuous domain, while also meeting the requirements for high-resolution detection, in the absence of online communication interaction.
A continuous transmission and reception signal model for a networked radar is constructed. Through a multi-agent partially observable Markov decision process, a composite reward function is designed. A multi-agent deep deterministic policy gradient algorithm based on a centralized training and distributed execution architecture is applied to achieve continuous frequency band resource scheduling for the radar, overcome the multi-agent credit allocation problem, and introduce a self-attention mechanism to adaptively extract time-varying mutual interference topology.
It enables radars to flexibly expand their operating bandwidth by utilizing fragmented spectrum gaps, ensuring high-range resolution detection. It significantly enhances the coordination and battlefield survivability of networked radars in strong electromagnetic suppression environments, and effectively resolves the strategic coupling and environmental non-stationarity among multiple agents.
Smart Images

Figure CN122085227A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of radar anti-jamming technology, and in particular to a network radar decision-making method based on an improved MADDPG that balances resolution and anti-jamming considerations. Background Technology
[0002] With the rapid evolution of electronic warfare technology, the electromagnetic environment of the modern battlefield is characterized by extreme complexity and high dynamism. On the one hand, enemy multi-point coordinated broadband frequency-sweeping jammers can rapidly cover and segment radar operating frequency bands, resulting in a fragmented and discontinuous distribution of the available net spectrum for radar. On the other hand, radar systems are rapidly shifting towards multi-station networked collaborative detection. While this improves detection performance, as the network scale expands, severe co-channel interference is highly likely to occur between stations within a limited frequency domain. Therefore, how to dynamically embed radar signals into the increasingly fragmented spectrum "gaps," while simultaneously avoiding external broadband interference and eliminating internal adjacent-channel interference, and obtaining as much continuous operating bandwidth as possible to ensure high-range resolution detection performance of the radar, has become a key challenge that urgently needs to be addressed to improve the overall combat effectiveness of networked radars.
[0003] To address the aforementioned challenges, data-driven reinforcement learning methods, due to their independence from prior environmental models and strong adaptive decision-making capabilities, have gradually become a research hotspot for solving the problem of intelligent radar anti-jamming. Existing research has made some progress in radar anti-jamming using reinforcement learning. In existing technologies, a frequency-agile strategy based on decentralized Q-learning has been proposed to address the dual problems of external interference and internal mutual interference faced by multi-radar systems. This strategy models the multi-radar collaborative process as a generalized Markov decision process, achieving joint suppression of interference and mutual interference through cooperation among agents. However, this method, based on the traditional Q-learning framework, restricts its decision space to a pre-defined set of discrete frequency points, limiting the radar to a fixed operating bandwidth. This discretization not only prevents the radar from fully utilizing the fragmented continuous spectrum resources between interfering frequency bands but also prevents the radar from improving range resolution by dynamically expanding signal bandwidth. Furthermore, to address the spectrum congestion problem, existing technologies have explored using deep Q-networks to control the radar's center frequency and bandwidth. The aim is to avoid communication interference through intelligent decision-making while maximizing the use of remaining spectrum gaps to maintain sufficient effective bandwidth, thereby effectively balancing target detection probability and range resolution in complex electromagnetic environments. Although this study verifies the importance of bandwidth adjustment in anti-jamming, it only involves the case of a single radar and cannot directly meet the needs of multiple radars. At the same time, the DQN algorithm is essentially still a value iteration method for handling discrete actions, and often faces the curse of dimensionality or quantization error when facing high-dimensional continuous action spaces.
[0004] Currently, in multi-radar networking scenarios, there is still a lack of effective solutions for simultaneously addressing broadband frequency sweep interference avoidance and adjacent frequency interference reduction through collaborative decision-making in the continuous domain, while also meeting the requirements for high-resolution detection, under conditions where online communication interaction is lacking. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a networked radar decision-making method based on an improved MADDPG that balances resolution and anti-interference capabilities, further enhancing radar survivability and detection in adversarial environments.
[0006] This invention provides a network radar decision-making method based on an improved MADDPG, which balances resolution and interference resistance, comprising the following steps:
[0007] Step 1: Construct an adversarial scenario between a networked radar and a multi-point frequency-sweeping jammer, and establish a radar continuous transmission signal model and a received signal model that support dynamic adjustment of frequency band boundaries;
[0008] Step 1.1: Configure the networked radar system to include A group of radars, operating in a single-transmit, single-receiver mode, are spatially distributed and maintain pulse synchronization in time. This radar set is denoted as _____. The scenario involves a target aircraft in a simulated combat environment, carrying a multi-point frequency-sweeping self-defense jammer. This jammer emits signals during each pulse repetition interval. Each bandwidth is Interference subband;
[0009] In a networked radar system, each radar transmits a continuously linearly frequency-modulated signal. The agility action of radar i on the m-th pulse is parameterized by a continuous vector, directly specifying the continuous starting frequency of the spectrum. With continuous termination frequency ;
[0010] Then the continuous transmission signal of radar i in the m-th pulse The model is based on a continuous radar signal transmission model, as detailed below:
[0011] ;
[0012] in, For peak transmission power, The pulse width. It is a rectangular pulse function. For frequency modulation slope, The imaginary unit, For time, center frequency With instantaneous operating bandwidth All are derived quantities uniquely determined by the decision variables:
[0013] ;
[0014] ;
[0015] Furthermore, the selected frequency band must meet the minimum bandwidth of the hardware system. With maximum bandwidth Constraints, namely ;
[0016] Step 1.3: Place the radar In the The total received signal of each pulse The model is based on the received signal model, as follows:
[0017] ;
[0018] in, For the target echo signal, For external multi-point frequency sweep interference signals, This is an internal system interference signal. It is additive white Gaussian noise;
[0019] Step 1.4: Calculate the frequency domain overlap ratio between the radar receiver and each radiation source, solve for the effective suppression power entering the detection channel, and calculate the signal-to-interference-plus-noise ratio (SNR) of the underlying physical reference output. Specifically:
[0020] Let there be any interference source or mutual interference source. The instantaneous radiation frequency band is The transmission bandwidth is Then the source of interference or mutual interference With radar Frequency domain intersection length The calculation is as follows:
[0021] ;
[0022] Define the spectral overlap coefficient This is the ratio of the frequency domain intersection length to the transmission bandwidth.
[0023] ;
[0024] Using the aforementioned spectral overlap coefficient to analyze the received signal The power of each component is modulated to obtain the actual power entering the radar. Effective suppression interference power of the detection channel Interference power with effective system :
[0025] ;
[0026] ;
[0027] in, For interference sources to radar Absolute suppression power For the spectral overlap coefficient of the interference source, For mutual interference radar radar absolute mutual interference power, For mutual interference sources to radar Absolute suppression power and spectral overlap coefficient, To remove radar Interference radar index, To suppress interference source indexes.
[0028] radar The received signal is subjected to pulse compression processing, and the obtained pulse compression gain is... It is proportional to the product of the radar's instantaneous operating bandwidth and pulse width, i.e. Therefore, radar In the The output signal-to-interference-plus-noise ratio of each pulse for:
[0029] ;
[0030] in, For radar Target echo signal power, For radar Received system thermal noise power.
[0031] Step 2: Analyze the frequency band overlap distribution. Based on the overlap between the radar operating frequency band and external interference and internal mutual interference, calculate the contaminated dirty bandwidth and the uncontaminated clean bandwidth respectively, and construct an optimized model that takes into account both detection resolution and anti-interference performance.
[0032] Step 2.1: For continuous frequency band resource scheduling, the degree of contamination is quantified by calculating the geometric overlap length of the signal frequency band in the frequency domain; assuming radar... In the The instantaneous operating frequency band of each pulse is The external multi-point broadband frequency sweep jammer transmitted the first One source of interference is Other friendly radars within the network The operating frequency band is ,in ;
[0033] Step 2.2: Define the radar The sum of the intersections of the operating frequency band with all external interference subbands and the operating frequency bands of other radars within the network constitutes the contaminated dirty bandwidth. The calculation formula is: ,in This represents the total number of interfering subbands.
[0034] Step 2.3: Define the radar Uncontaminated clean bandwidth Instantaneous operating bandwidth The remaining portion after deducting dirty bandwidth is:
[0035] ;
[0036] Step 2.4: With the joint objective of maximizing the radar's continuous clean bandwidth and minimizing the contaminated dirty bandwidth, construct a utility optimization model that balances detection resolution and spectral purity, as follows:
[0037] ;
[0038] in, Indicates the total number of pulses. This indicates the number of radar nodes in a networked radar system. and These represent the lower and upper limits of the available operating frequency bands supported by the radar hardware system, respectively.
[0039] Step 3: Model the collaborative scheduling process of the networked radar as a multi-agent partially observable Markov decision process, directly define the radar action as the start and end frequencies of the continuous frequency band, and design a composite reward function that balances bandwidth expansion and conflict avoidance based on the optimization model.
[0040] Step 3.1: Model the cooperative scheduling problem of the networked radar system as a partially observable Markov decision process, and define the global state space. With local observation space The global state space includes local observations from all radars and the historical mutual interference ratios of each radar. The local observation space includes the interference source perception characteristics of a single radar, the action at the previous moment, and the reward at the previous moment.
[0041] Define a continuous joint action space :radar At any moment action Defined as the available radio frequency passband in the system Selected continuous start frequency With termination frequency A two-dimensional vector, namely:
[0042] ;
[0043] Step 3.3: Based on the optimization model, design a composite reward function, specifically by radar. At any moment The single-step basic utility gained is derived from positive clean bandwidth benefits. Penalty for negative dirty bandwidth conflicts constitute:
[0044] ;
[0045] ;
[0046] in, Dirty bandwidth penalty weight, This represents the radar's instantaneous total operating bandwidth.
[0047] Step 3.4: Introduce a global fairness mechanism and define the radar. At any moment The final partial reward obtained for: ;
[0048] in, , For the fairness term, the weighting coefficient is... This is the soft minimum of the current base utility value for all radar nodes.
[0049] Step 4: Apply the multi-agent deep deterministic policy gradient algorithm based on centralized training and distributed execution architecture to select actions and evaluate values, solve the Markov decision process, and obtain the optimal continuous frequency band resource collaborative scheduling strategy for each radar.
[0050] Step 4.1: For each network radar Construct a policy network based on a centralized training-distributed execution architecture With value network And initialize the corresponding target network parameters and experience replay pool. Among them, radar Local observations as policy networks The input is processed through three fully connected layers to obtain continuous frequency band actions as output values; during the intensive training phase, the value network... A global action value assessment is performed by acquiring local observations from all radars. The action value function satisfies the Bellman expectation equation, which is specifically expressed as follows:
[0051] ;
[0052] in, Let t be the global state. This represents the global state at time t+1. Discount factor; Let t represent the actions of each radar. The actions of each radar at time t+1;
[0053] Step 4.2: In the value network The feature extraction module is constructed in the middle, specifically: the radar Local observation at time t ,action and historical interference ratio Input multilayer perceptron After passing through two fully connected layers, they are mapped into independent high-dimensional semantic vectors. :
[0054] ;
[0055] Step 4.3: Aggregate the node feature sequences The multi-head self-attention layer is fed into the system for deep interaction. Based on the dot product similarity calculation between the query and the key, features of radar nodes that are similar to or pose a severe threat of mutual interference in the current frequency band are adaptively extracted.
[0056] ;
[0057] in, This represents a multi-head attention mechanism, using feature sequences that fuse attention weights. After being flattened, the input is evaluated from the head, and the output is a centralized motion value. ;
[0058] Step 4.4: After the networked radar system executes joint continuous frequency band actions, it obtains the reward and next state from the environmental feedback and stores them in the experience replay pool. ;from Randomly sampled batches of data, minimizing the time-series difference error loss. Update value network parameters:
[0059] ;
[0060] ;
[0061] in, For radar Value network parameters; An operator for finding mathematical expectation; For radar The target value network function, The action at time t+1 is the output of the target policy network; This is the end-of-round flag; it is 1 when the round ends and 0 otherwise.
[0062] Step 4.5: Use the updated value network to guide the policy network update, adjusting the policy parameters along the direction of increasing value using deterministic policy gradients. :
[0063] ;
[0064] in, For radar The strategy network parameters, Let the policy objective function be... For experience replay pool, For radar The policy network function, Indicates the network parameters related to the policy. gradient operator, Indicates about actions The gradient operator;
[0065] Step 4.6: Use a soft update mechanism to periodically update the parameters of the target policy network and the target value network to ensure the stability of the target evaluation value during training;
[0066] After the training phase is completed and the parameters of both the policy network and the value network converge, the value network is removed. During the online execution phase, each network radar node retains and runs only its own policy network. It directly outputs the start and end frequencies of continuous frequency bands based on real-time local observations, thereby achieving optimal continuous frequency band collaborative scheduling under conditions without online communication interaction.
[0067] The beneficial effects of adopting the above technical solution are as follows:
[0068] This invention provides a networked radar decision-making method based on an improved MADDPG that balances resolution and anti-interference capabilities. Compared with existing technologies, this invention departs from the traditional discrete frequency hopping mode, expanding the radar action space to a continuous frequency band boundary. This allows the radar to flexibly utilize fragmented spectrum gaps to elastically expand its operating bandwidth, effectively ensuring high-range resolution detection requirements. The constructed reward mechanism provides a stable gradient for strategy exploration in the high-dimensional continuous action space, overcoming the multi-agent credit allocation problem and achieving a dynamic trade-off between detection accuracy and anti-interference survivability. Furthermore, the introduction of a self-attention mechanism adaptively extracts time-varying mutual interference topology, effectively mitigating strategy coupling among multiple agents and environmental non-stationarity. Finally, with an online cooperative strategy that has extremely low communication dependence, the networked radar can simultaneously counter external multi-point frequency sweeping interference and suppress internal mutual interference, significantly enhancing the cooperativeness and battlefield survivability of the networked radar in strong electromagnetic suppression environments. Attached Figure Description
[0069] Figure 1 This is a flowchart of the network radar decision-making method of the present invention, which takes into account both resolution and anti-interference.
[0070] Figure 2 This is a schematic diagram of networked radar cooperative detection in an embodiment of the present invention.
[0071] Figure 3 This is a schematic diagram illustrating the interaction between the algorithm and the environment in an embodiment of the present invention.
[0072] Figure 4 These are the reward convergence curves of each radar and the total reward convergence curve of the networked radar under the simulation experiment in the embodiments of the present invention; wherein, (a) is the reward convergence curve of each radar, and (b) is the total reward convergence curve of the networked radar.
[0073] Figure 5 The graph shows the variation of the average output signal-to-interference-plus-noise ratio (SINJR) of each radar under the simulation experiment in this embodiment of the invention with the number of iterations.
[0074] Figure 6 This is a diagram showing the frequency band selection changes of each radar in a round of evaluation during the simulation experiment in this embodiment of the invention.
[0075] Figure 7 The figures show the changes in the interference bandwidth ratio and mutual interference bandwidth ratio of each radar under the simulation experiment in the embodiments of the present invention; wherein, (a) is the interference bandwidth ratio of each radar, and (b) is the change in the mutual interference bandwidth ratio of each radar. Detailed Implementation
[0076] The specific implementation methods of this application will be further described in detail below with reference to the accompanying drawings and embodiments.
[0077] Example 1:
[0078] This invention provides a network radar decision-making method based on an improved MADDPG, which balances resolution and interference resistance. Figure 1 As shown, it includes the following steps:
[0079] Step 1: Construct an adversarial scenario between a networked radar and a multi-point frequency-sweeping jammer, and establish a radar continuous transmission signal model and a received signal model that support dynamic adjustment of frequency band boundaries;
[0080] Step 1.1: Configure the networked radar system to include The radars operate collaboratively in a single-transmit, single-receiver mode, distributed spatially and maintaining pulse synchronization in time, as shown in the diagram. Figure 2 As shown, let the radar set be... The scenario involves a target aircraft in a simulated combat environment, carrying a multi-point frequency-sweeping self-defense jammer. This jammer emits signals during each pulse repetition interval. Each bandwidth is Interference subband;
[0081] In a networked radar system, each radar transmits a continuously linearly frequency-modulated signal. The agility action of radar i on the m-th pulse is parameterized by a continuous vector, directly specifying the continuous starting frequency of the spectrum. With continuous termination frequency ;
[0082] Then the continuous transmission signal of radar i in the m-th pulse The model is based on a continuous radar signal transmission model, as detailed below:
[0083] ;
[0084] in, For peak transmission power, The pulse width. It is a rectangular pulse function. For frequency modulation slope, The imaginary unit, For time, center frequency With instantaneous operating bandwidth All are derived quantities uniquely determined by the decision variables:
[0085] ;
[0086] ;
[0087] Furthermore, the selected frequency band must meet the minimum bandwidth of the hardware system. With maximum bandwidth Constraints, namely ;
[0088] Step 1.3: Place the radar In the The total received signal of each pulse The model is based on the received signal model, as follows:
[0089] ;
[0090] in, For the target echo signal, For external multi-point frequency sweep interference signals, This is an internal system interference signal. It is additive white Gaussian noise;
[0091] Step 1.4: Calculate the frequency domain overlap ratio between the radar receiver and each radiation source, solve for the effective suppression power entering the detection channel, and calculate the signal-to-interference-plus-noise ratio (SNR) of the underlying physical reference output. Specifically:
[0092] The RF front-end bandpass filter of a radar receiver determines that only frequencies falling within its instantaneous operating frequency band are allowed. The signal energy within can enter the subsequent processing channel;
[0093] Let there be any interference source or mutual interference source. The instantaneous radiation frequency band is The transmission bandwidth is Then the source of interference or mutual interference With radar Frequency domain intersection length The calculation is as follows:
[0094] ;
[0095] Define the spectral overlap coefficient This is the ratio of the frequency domain intersection length to the transmission bandwidth.
[0096] ;
[0097] Using the aforementioned spectral overlap coefficient to analyze the received signal The power of each component is modulated to obtain the actual power entering the radar. Effective suppression interference power of the detection channel Interference power with effective system :
[0098] ;
[0099] ;
[0100] in, For interference sources to radar Absolute suppression power For the spectral overlap coefficient of the interference source, For mutual interference radar radar absolute mutual interference power, For mutual interference sources to radar Absolute suppression power and spectral overlap coefficient, To remove radar Interference radar index, To suppress interference source indexes.
[0101] radar The received signal is subjected to pulse compression processing, and the obtained pulse compression gain is... It is proportional to the product of the radar's instantaneous operating bandwidth and pulse width, i.e. Therefore, radar In the The output signal-to-interference-plus-noise ratio of each pulse for:
[0102] ;
[0103] in, For radar Target echo signal power, For radar Received system thermal noise power.
[0104] The output signal-to-interference-plus-noise ratio characterizes the radar's actual detection and anti-jamming capabilities under the consideration of physical frequency band overlap masking effects, and will be used as the underlying evaluation benchmark.
[0105] Step 2: Analyze the frequency band overlap distribution. Based on the overlap between the radar operating frequency band and external interference and internal mutual interference, calculate the contaminated dirty bandwidth and the uncontaminated clean bandwidth respectively, and construct an optimized model that takes into account both detection resolution and anti-interference performance.
[0106] Step 2.1: For continuous frequency band resource scheduling, the degree of contamination is quantified by calculating the geometric overlap length of the signal frequency band in the frequency domain; assuming radar... In the The instantaneous operating frequency band of each pulse is The external multi-point broadband frequency sweep jammer transmitted the first One source of interference is Other friendly radars within the network The operating frequency band is ,in ;
[0107] Step 2.2: Define the radar The sum of the intersections of the operating frequency band with all external interference subbands and the operating frequency bands of other radars within the network constitutes the contaminated dirty bandwidth. Its physical meaning is the bandwidth of the frequency band that enters the radar receiver and generates effective electromagnetic suppression, and its calculation formula is: ,in This represents the total number of interfering subbands.
[0108] The specific physical implementation of the intersection operation is as follows: for any two frequency bands and Its overlap bandwidth length function The calculations are as follows: ;
[0109] Step 2.3: Define the radar Uncontaminated clean bandwidth Instantaneous operating bandwidth The remaining portion after deducting the dirty bandwidth determines the actual detection resolution that the radar can achieve without being suppressed, namely:
[0110] ;
[0111] Step 2.4: With the joint objective of maximizing the radar's continuous clean bandwidth and minimizing the contaminated dirty bandwidth, construct a utility optimization model that balances detection resolution and spectral purity, as follows:
[0112] ;
[0113] in, Indicates the total number of pulses. This indicates the number of radar nodes in a networked radar system. and These represent the lower and upper limits of the available operating frequency bands supported by the radar hardware system, respectively.
[0114] Step 3: Model the collaborative scheduling process of the networked radar as a multi-agent partially observable Markov decision process, directly define the radar action as the start and end frequencies of the continuous frequency band, and design a composite reward function that balances bandwidth expansion and conflict avoidance based on the optimization model.
[0115] Step 3.1: Model the cooperative scheduling problem of the networked radar system as a partially observable Markov decision process, and define the global state space. With local observation space The global state space includes local observations from all radars and the historical mutual interference ratios of each radar. The value network is only visible during the centralized training phase; the local observation space includes the interference source perception characteristics of a single radar, the action at the previous time step, and the reward at the previous time step, which are used by each radar policy network to make independent decisions during the distributed execution phase; the algorithm's interaction with the environment is as follows: Figure 3 For radar i, For local observation, For each action, i ∈ [1, N];
[0116] Define a continuous joint action space :radar At any moment action Defined as the available radio frequency passband in the system Selected continuous start frequency With termination frequency A two-dimensional vector, namely:
[0117] ;
[0118] Step 3.3: Based on the optimization model, directly utilize "clean bandwidth". With "dirty bandwidth" Design a composite reward function, specifically by radar. At any moment The single-step basic utility gained is derived from positive clean bandwidth benefits. Penalty for negative dirty bandwidth conflicts constitute:
[0119] ;
[0120] ;
[0121] in, Dirty bandwidth penalty weight, This represents the radar's instantaneous total operating bandwidth.
[0122] Step 3.4: To overcome the problem of uneven system resource allocation caused by the completely self-serving exploration during multi-agent frequency band contention, a global fairness term mechanism is introduced based on the aforementioned basic utility, defining the radar... At any moment The final partial reward obtained for: ;
[0123] in, , This is the weighting coefficient for the fairness item, used to adjust the degree of inclination between individual gains and system fairness; This is the soft minimum of the current base utility value for all radar nodes.
[0124] Step 4: Apply the multi-agent deep deterministic policy gradient algorithm based on centralized training and distributed execution architecture to select actions and evaluate values, solve the Markov decision process, and obtain the optimal continuous frequency band resource collaborative scheduling strategy for each radar.
[0125] Step 4.1: For each network radar Construct a policy network based on a centralized training-distributed execution architecture With value network And initialize the corresponding target network parameters and experience replay pool. Among them, radar Local observations as policy networks The input is processed through three fully connected layers to obtain continuous frequency band actions as output values; during the intensive training phase, the value network... A global action value assessment is performed by acquiring local observations from all radars. The action value function satisfies the Bellman expectation equation, which is specifically expressed as follows:
[0126] ;
[0127] in, Let t be the global state. This represents the global state at time t+1. Discount factor; Let t represent the actions of each radar. The actions of each radar at time t+1;
[0128] Step 4.2: In the value network The feature extraction module is constructed in the middle, specifically: the radar Local observation at time t ,action and historical interference ratio Input multilayer perceptron After passing through two fully connected layers, they are mapped into independent high-dimensional semantic vectors. :
[0129] ;
[0130] Step 4.3: Aggregate the node feature sequences The multi-head self-attention layer is fed into the system for deep interaction. Based on the dot product similarity calculation between the query and the key, features of radar nodes that are similar to or pose a severe threat of mutual interference in the current frequency band are adaptively extracted.
[0131] ;
[0132] in, This represents a multi-head attention mechanism, using feature sequences that fuse attention weights. After being flattened, the input is evaluated from the head, and the output is a centralized motion value. ;
[0133] Step 4.4: After the networked radar system executes joint continuous frequency band actions, it obtains the reward and next state from the environmental feedback and stores them in the experience replay pool. ;from Randomly sampled batches of data, minimizing the time-series difference error loss. Update value network parameters:
[0134] ;
[0135] ;
[0136] in, For radar Value network parameters; An operator for finding mathematical expectation; For radar The target value network function, The action at time t+1 is the output of the target policy network; This is the end-of-round flag; it is 1 when the round ends and 0 otherwise.
[0137] Step 4.5: Use the updated value network to guide the policy network update, adjusting the policy parameters along the direction of increasing value using deterministic policy gradients. :
[0138] ;
[0139] in, For radar The strategy network parameters, Let the policy objective function be... For experience replay pool, For radar The policy network function, Indicates the network parameters related to the policy. gradient operator, Indicates about actions The gradient operator;
[0140] Step 4.6: Use a soft update mechanism to periodically update the parameters of the target policy network and the target value network to ensure the stability of the target evaluation value during training;
[0141] After the training phase is completed and the parameters of both the policy network and the value network converge, the value network is removed. During the online execution phase, each network radar node retains and runs only its own policy network. It directly outputs the start and end frequencies of continuous frequency bands based on real-time local observations, thereby achieving optimal continuous frequency band collaborative scheduling under conditions without online communication interaction.
[0142] Example 2:
[0143] The present invention also provides another embodiment, which verifies and analyzes the method of the present invention through simulation:
[0144] During the collaborative detection process of networked radars, each radar node adjusts the start and end frequencies of its transmitted signals to avoid multi-point frequency sweep interference, suppress internal co-frequency interference, and optimize the resolution of the transmitted signals. Each radar transmits a continuous linear frequency modulated (LFM) signal, which is scattered by the target to form an echo and then received by the radar. Based on the spectrum pollution situation, frequency band resources are dynamically and adaptively scheduled.
[0145] In the cooperative detection scenario, the number of networked radar nodes is set to 3. A target aircraft carrying a multi-point broadband sweeping self-defense jammer exists in the battlefield environment. The radar coordinates in the Cartesian coordinate system are [-3.2km, 0km, 0.5km], [14.0km, 18.0km, 1.5km], and [-25.0km, 30.0km, 0.9km], respectively. The target coordinates are [-110.0km, 90.0km, 8.0km], and the target RCS is 30m². The total system radio frequency bandwidth is set to 4.0GHz to 4.6GHz, the radar pulse width is 40μs, the pulse repetition interval (PRI) is 400μs, and the dynamic adjustment boundary of the radar instantaneous operating bandwidth is set to a minimum of 20MHz and a maximum of 100MHz. The jammer employs a "multi-point random jump" frequency sweeping strategy, randomly transmitting three jamming sub-bands within each PRI (Primary Principle), each sub-band having a bandwidth of 100MHz and a frequency step interval set to 100MHz. In the MADDPG algorithm training hyperparameter settings, the maximum number of training epochs is 1000, the maximum number of steps per epoch is 100, and the learning rates for the Actor network and Critic network are set to 3×10⁻⁶ respectively. -4 and 5×10 -4 The discount factor γ is set to 0.8, the target network soft update coefficient τ is set to 0.005, and the target network update frequency is 4. In the reward function parameter settings, λ... d A value of 1 is used to balance bandwidth utilization and collision penalty; an excessively large λ... d This can lead to lower bandwidth utilization. If the system has higher requirements for frequency band conflict suppression, λ can be appropriately increased. d The value of β is set to 0.3, so that the fairness term only participates in the optimization as an auxiliary correction term, avoiding its excessive coverage of the basic reward term.
[0146] Figure 4 (a) is a graph showing the reward convergence curves for each radar. Figure 4 (b) is a graph showing the overall reward convergence curve of the networked radars, which includes the original reward trajectory of each radar and its moving average curve. Figure 4In (b), the team reward refers to the sum of the rewards for each radar, i.e., the total reward of the networked radar system. As shown in the figure, in the early training phase (approximately 0 to 200 iterations), due to the blind exploration phase of each radar and the randomness of frequency band selection leading to severe mutual interference, the discounted round rewards for each radar node were low and fluctuated wildly, even showing negative values. With the increase in the number of iterations, the reward curve showed a steady upward trend. After approximately 600 iterations, the curve gradually stabilized, and the individual rewards for radar 1, radar 2, and radar 3 converged and were closely distributed between 8 and 10, while the total reward of the networked radar system also stabilized at around 28. It can be seen that after iterative training, the agent can learn an effective continuous frequency band cooperative scheduling strategy, spontaneously completing frequency band orthogonal avoidance based on the current state, effectively achieving optimal convergence of high-precision detection and anti-interference / mutual interference comprehensive performance.
[0147] Figure 5 This figure shows the variation of the average output signal-to-interference-plus-noise ratio (SINJR) of each radar with the number of iterations in the simulation experiment of this embodiment. As can be seen from the figure, in the early training stage (approximately 0 to 200 iterations), due to the agent's blind exploration phase, the frequency band selection is highly random, leading to severe overlap between the operating frequency bands of each radar and with the mechanically scanned frequency bands of the jamming radar. The average SINJR of each radar is at a low level and exhibits significant oscillations. With the increase in the number of iterations, the average SINJR curve of each radar shows a steady upward trend. After approximately 600 iterations, the curve gradually stabilizes, and the average SINJR of radars 1, 2, and 3 converge and are closely distributed between 13dB and 14dB. It can be seen that after iterative training, the agent can learn an effective continuous frequency band cooperative scheduling strategy, significantly improving the target echo quality after pulse compression from the underlying energy dimension, verifying the effectiveness of this invention in improving the physical detection efficiency and real anti-jamming gain of networked radars.
[0148] Figure 6 This figure shows the frequency band selection changes of each radar during a round of evaluation in a simulation experiment of this invention. The light gray rectangular shaded area represents the dynamic frequency sweeping jamming band of the multi-point broadband frequency sweeping jammer. The broken lines with different line types and error bars represent the actual center frequencies and bandwidth ranges of radars 1, 2, and 3, respectively. As shown in the figure, under the condition of no online communication during the execution phase, all three radars accurately avoided the dynamically changing gray jamming area, and there was almost no overlap between the transmission frequency bands of each radar. It can be seen that after iterative training, the agent can learn the frequency sweeping rules of the jammer and the frequency usage habits of its teammates, intuitively confirming that the CTDE architecture of this invention can guide multiple radars to implicitly form a highly tacit spectrum orthogonal avoidance and bandwidth elastic scaling mechanism in the absence of communication.
[0149] Figure 7This is a graph showing the changes in the percentage of interference bandwidth and the percentage of mutual interference bandwidth for each radar under the simulation experiment in this embodiment of the invention. Figure 7 (a) is a diagram showing the percentage of interference bandwidth for each radar. Figure 7 (b) is a graph showing the change in the cross-interference bandwidth ratio of each radar. As can be seen from the graph, in the early stages of training, due to the blind selection of frequencies, the cross-interference bandwidth ratio and cross-interference bandwidth ratio of each radar were at relatively high levels. The cross-interference bandwidth ratio reached a maximum of approximately 0.45, while the cross-interference bandwidth ratio even peaked at over 0.6 in the very early stages. With the increase in iterations, each radar learned avoidance strategies under the guidance of the reward function, and both curves showed a significant downward trend. After approximately 600 iterations, the cross-interference bandwidth ratio of each radar node continued to decrease, eventually converging stably between 0.1 and 0.15; while the cross-interference bandwidth ratio of all three radars converged to an extremely low state close to 0. It can be seen that after iterative training, the method of this invention successfully decoupled and dual-suppressed external suppression threats and internal resource conflicts at the macro-statistical level, demonstrating the extremely high stability of the designed composite reward function in solving the multi-agent continuous spectrum pollution problem.
[0150] In summary, the method of this invention constructs the process of networked radars coordinating to resist external suppression interference and suppress internal mutual interference in complex electromagnetic environments as a multi-agent partially observable Markov decision process. This invention treats local perception containing interference source characteristics as observation, the start and end frequencies of continuous frequency bands as actions, and the dual smoothing effect that balances detection resolution and spectral purity as the reward function. Relying on a centralized training-distributed execution architecture and a value evaluation network integrating multi-head self-attention, each radar node can independently complete intelligent collaborative scheduling of frequency band resources based solely on local historical observations, without online communication interaction. This method guides multiple radars to spontaneously form a tacit spectrum orthogonal avoidance and bandwidth elastic scaling strategy, thereby almost completely eliminating internal co-frequency mutual interference while resisting broadband frequency sweeping interference, effectively achieving a comprehensive improvement in high-precision detection and high anti-interference survivability.
[0151] Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a computer program product.
[0152] The various embodiments in this application are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0153] The scope of protection of this application is not limited to the embodiments described above. Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from the scope and spirit of this disclosure. If such modifications and variations fall within the scope of the methods disclosed herein and their equivalents, then the intent of this disclosure also includes such modifications and variations.
Claims
1. A networking radar decision-making method that balances resolution and anti-jamming based on improved MADDPG, characterized in that, Includes the following steps: Step 1: Construct an adversarial scenario between a networked radar and a multi-point frequency-sweeping jammer, and establish a radar continuous transmission signal model and a received signal model that support dynamic adjustment of frequency band boundaries; Step 2: Analyze the frequency band overlap distribution. Based on the overlap between the radar operating frequency band and external interference and internal mutual interference, calculate the contaminated dirty bandwidth and the uncontaminated clean bandwidth respectively, and construct an optimized model that takes into account both detection resolution and anti-interference performance. Step 3: Model the collaborative scheduling process of the networked radar as a multi-agent partially observable Markov decision process, directly define the radar action as the start and end frequencies of the continuous frequency band, and design a composite reward function that balances bandwidth expansion and conflict avoidance based on the optimization model. Step 4: Apply the multi-agent deep deterministic policy gradient algorithm based on centralized training and distributed execution architecture to select actions and evaluate value, solve the Markov decision process, and obtain the optimal continuous frequency band resource collaborative scheduling strategy for each radar.
2. The networking radar resolution and anti-jamming decision method based on improved MADDPG according to claim 1, characterized in that, Step 1 specifically includes the following steps: Step 1.1: Set up a netted radar system containing M radars, working in monopulse monostatic mode, distributed in space and keeping pulse synchronization in time, record the radar set as ; Set up an environment with a target aircraft, carrying a multi-point swept self-defense jammer, which transmits interference sub-bands with a bandwidth of A set of networking radar system in each radar transmitting continuous linear frequency modulation signal, radar i in the mth pulse of the agile action by continuous vector parameterization, directly designated the continuous start frequency of the spectrum and continuous end frequency The radar i in the mth pulse of the continuous transmission signal Modeling as a radar continuous transmission signal model, specifically as follows: ; where is the peak transmit power, is the pulse width, is the rectangular pulse function, is the frequency modulation slope, is the imaginary unit, is the time instant, the center frequency is the instantaneous operating bandwidth are both derived quantities uniquely determined by the decision variables: ; ; and the selected frequency band needs to meet the minimum bandwidth of the hardware system with the constraint of the maximum bandwidth , i.e. ; Step 1.3: The radar receives the total received signal of the first pulse is modeled as a received signal model, in particular as follows: ; wherein, is a target echo signal, is an external multi-point swept interference signal, is an internal system mutual interference signal, is additive white Gaussian noise; Step 1.4: Calculate the frequency domain overlap ratio between the radar receiver and each radiation source, solve for the effective suppression power entering the detection channel, and calculate the signal-to-interference-plus-noise ratio of the underlying physical reference output.
3. The networked radar resolution and anti-jamming decision method based on improved MADDPG according to claim 2, characterized in that, Step 2 specifically includes the following steps: Step 2.1: For continuous frequency band resource scheduling, the pollution degree is quantified by calculating the geometric overlap length of the signal frequency band in the frequency domain; let the radar work in the frequency band of , the instantaneous operating frequency band of the th pulse is , the th interference source emitted by the external multi-point broadband sweep jammer is , and the operating frequency band of other friendly radars in the network is , wherein ; Step 2.2: Define the radar The sum of the intersection of the operating frequency band of the radar and all external interference sub-bands and the operating frequency bands of other radars in the network is the contaminated dirty bandwidth The calculation formula is: Wherein is the total number of interference sub-bands; Step 2.3: Define Radar Uncontaminated clean bandwidth For instantaneous operating bandwidth The remainder after deducting the dirty bandwidth, i.e.: ; Step 2.4: With the joint objective of maximizing the radar's continuous clean bandwidth and minimizing the contaminated dirty bandwidth, construct a utility optimization model that balances detection resolution and spectral purity, as follows: ; wherein, denotes the total number of pulses, denotes the number of radar nodes in the networked radar system, and denote the lower and upper frequency of the available operating frequency band supported by the radar hardware system, respectively.
4. The networking radar resolution and anti-jamming decision method based on improved MADDPG according to claim 3, characterized in that, Step 3 specifically includes the following steps: Step 3.1: Model the cooperative scheduling problem of the networked radar system as a partially observable Markov decision process, define the global state space and the local observation space ; wherein the global state space contains the local observations of all radars and the historical mutual interference ratios of each radar , and the local observation space contains the interference source perception features of a single radar itself, the action at the last time and the reward at the last time; Defining continuous joint action space : radar At time The action is defined as a two-dimensional vector of a selected continuous start frequency and end frequency within the system's available radio frequency passband , i.e.: ; Step 3.3: Design a composite reward function based on the optimization model, specifically by Radar At time The single-step base utility obtained is composed of a positive clean bandwidth reward and a negative dirty bandwidth conflict penalty : ; ; wherein is a dirty bandwidth penalty weight, is a radar instantaneous total operating bandwidth; Step 3.4: Introduce a global fairness item mechanism, define radar At time Final acquired local reward For: ; wherein, , is a fairness weight coefficient, is a soft minimum of the current base utility values of all radar nodes.
5. The networking radar resolution and anti-jamming decision method based on improved MADDPG according to claim 4, characterized in that, Step 4 specifically includes the following steps: Step 4.1: for each networking radar Constructing a policy network based on a centralized training-distributed execution architecture With a value network , and initializing the corresponding target network parameters and experience replay pool ; wherein the local observation of the radar is taken as the input of the policy network , and after passing through three fully connected layers, a continuous frequency band action is obtained as the output value; in the centralized training stage, the value network obtains the local observation of all radars to perform global action value evaluation, and the action value function satisfies the Bellman expectation equation, which is specifically expressed as: ; wherein, is the global state at time t; is the global state at time t+1; is the discount factor; is the action of each radar at time t; is the action of each radar at time t+1; Step 4.2: Constructing feature extraction module in the value network , specifically: inputting local observations , actions , and historical mutual interference ratios at time t into a multi-layer perception , and mapping them into independent high-dimensional semantic vectors through two fully connected layers : ; Step 4.3: Aggregate the node feature sequences The multi-head self-attention layer is fed into the system for deep interaction. Based on the dot product similarity calculation between the query and the key, features of radar nodes that are similar to or pose a severe threat of mutual interference in the current frequency band are adaptively extracted. ; wherein, represents a multi-head attention mechanism, and the feature sequence fused with the attention weight input evaluation head after flattening, output centralized action value ; Step 4.4: After the networked radar system performs joint contiguous band action, the reward of the environment feedback and the next state are obtained and stored in the experience replay pool ; a batch of data is randomly sampled from the minimum time difference error loss update the value network parameters: ; ; wherein, is a radar value network parameter; is an operator for summing mathematical expectations; is a radar target value network function, is a t+1 time action output by the target policy network; is a round end flag, which takes a value of 1 when the round ends, and 0 otherwise; Step 4.5: Update the policy network guided by the updated value network, adjust the policy parameters in the direction of increasing value by deterministic policy gradient : ; wherein is a radar policy network parameter, is a policy objective function, is an experience replay pool, is a radar policy network function, denotes a gradient operator with respect to the policy network parameters denotes a gradient operator with respect to the action denotes a gradient operator with respect to the action denotes a gradient operator with respect to the action Step 4.6: Use a soft update mechanism to periodically update the parameters of the target policy network and the target value network to ensure the stability of the target evaluation value during training; After the training phase is completed and the parameters of both the policy network and the value network converge, the value network is removed. During the online execution phase, each network radar node retains and runs only its own policy network. It directly outputs the start and end frequencies of continuous frequency bands based on real-time local observations, thereby achieving optimal continuous frequency band collaborative scheduling under conditions without online communication interaction.
6. A network radar decision-making method based on improved MADDPG, considering both resolution and anti-interference, as described in claim 5, is characterized in that... Step 1.4 specifically involves: Let there be any interference source or mutual interference source. The instantaneous radiation frequency band is The transmission bandwidth is Then the source of interference or mutual interference With radar Frequency domain intersection length The calculation is as follows: ; Define the spectral overlap coefficient This is the ratio of the frequency domain intersection length to the transmission bandwidth. ; Using the aforementioned spectral overlap coefficient to analyze the received signal The power of each component is modulated to obtain the actual power entering the radar. Effective suppression interference power of the detection channel Interference power with effective system : ; ; in, For interference sources to radar Absolute suppression power For the spectral overlap coefficient of the interference source, For mutual interference radar radar absolute mutual interference power, For mutual interference sources to radar Absolute suppression power and spectral overlap coefficient, To remove radar Interference radar index, To suppress interference source index; radar The received signal is subjected to pulse compression processing, and the obtained pulse compression gain is... It is proportional to the product of the radar's instantaneous operating bandwidth and pulse width, i.e. Therefore, radar In the The output signal-to-interference-plus-noise ratio of each pulse for: ; in, For radar Target echo signal power, For radar Received system thermal noise power.