A reinforcement learning intelligent jamming method for networking frequency agility radar system

By optimizing the jamming strategy through reinforcement learning and combining frequency sweeping and response jamming, the problem of low efficiency of existing jamming systems in networked frequency-agile radar environments is solved, and a highly efficient intelligent jamming effect is achieved.

CN118444259BActive Publication Date: 2025-11-25UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410529772.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-29
Publication Date
2025-11-25
Estimated Expiration
2044-04-29

AI Technical Summary

Technical Problem

Existing jamming systems are ineffective against networked frequency-agile radars, especially in the absence of prior information. Traditional jamming strategies are inefficient and cannot effectively cope with the complex environment of multiple radars networked together.

Method used

A multifunctional jamming system is designed using reinforcement learning, combining frequency sweeping and response jamming. The jamming strategy is optimized through Markov decision process, and the optimal jamming strategy is trained using deep deterministic policy gradient algorithm, thereby achieving intelligent jamming of networked frequency-agile radar.

Benefits of technology

It improves the jamming efficiency against networked frequency-agile radars, enhances the adaptability and resource utilization efficiency of the jamming system, and significantly improves the strike capability in electronic warfare environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118444259B_ABST
    Figure CN118444259B_ABST
Patent Text Reader

Abstract

The application discloses a kind of network-oriented frequency agile radar system's reinforcement learning intelligent interference method, first establish the confrontation scene of multiple radar and jammer, and design a kind of pulse-by-pulse interference, mixed use point interference and sweep frequency interference interference system working model, then the time domain, energy domain resource allocation decision process of interference system when interfering with radar group is modeled as a Markov decision process, and the optimal policy of interference system is solved using deep deterministic policy gradient algorithm to train.This method of the application has strong adaptability and high resource utilization efficiency, compared with the existing interference method for frequency agile radar which depends on specific rules, can significantly enhance the interference efficiency of various configurations of networked frequency agile radar system, tap the task performance of jammer, while using deep reinforcement learning algorithm to optimize combined suppression jamming strategy, improve the strike capability of jammer in electronic countermeasure environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of electronic countermeasures technology, specifically relating to a reinforcement learning intelligent jamming method for networked frequency-agile radar systems. Background Technology

[0002] Modern radar anti-jamming mechanisms (frequency agility, waveform diversification, frequency hopping code division multiple access) are diverse, and anti-jamming methods are increasing, which puts forward new requirements for jamming attacks in electronic warfare. Jamming systems need to further develop corresponding jamming methods to target anti-jamming radars in the environment.

[0003] Among various anti-jamming methods, frequency agility is a widely used approach, capable of unpredictably changing the pulse carrier frequency over a broad bandwidth. Frequency-agile radar performs exceptionally well against jamming attacks, outperforming traditional radar. Through strategic design, frequency-agile radar exhibits strong anti-jamming capabilities, effectively resisting mainlobe interference. The various advantages of frequency-agile radar pose significant challenges to jamming systems, necessitating the development of new jamming strategies to address these issues.

[0004] Frequency scanning jamming is a conventional jamming method targeting frequency-agile radars. It operates by repeatedly jamming frequencies randomly or sequentially within a specific frequency band. However, with advancements in modern radar anti-jamming technology, the threat posed by frequency scanning jamming to radar systems is diminishing. A jamming mode called pulse-by-pulse jamming or reactive jamming has been designed to intercept each pulse emitted by the radar system and apply targeted jamming based on the intercepted information. However, it still requires a well-designed jamming strategy to achieve its ideal effect. To address this, an adaptive jamming strategy is proposed, which integrates various jamming modes and leverages the advantages of different jamming methods to significantly improve the effectiveness of the jamming system. However, designing a jamming strategy typically relies on sufficient prior information, which the jamming system often lacks.

[0005] Reinforcement learning can overcome this pain point of active anti-jamming because it does not rely on prior information, has strong environmental adaptability, and can effectively guide the generation of jamming strategies under conditions of scarce prior information. The paper "Jammingefficiency of land-based radars by the airborne jammers, 2018 22nd International Microwave and Radar Conference (MIKON)" proposes a jamming scenario using an airborne jammer against a primary radar, but it does not design a strategy to deal with radar frequency agility. The paper "A dynamic game strategy for radar screening pulse width allocation against jamming using reinforcement learning, IEEE Transactions on Aerospace and Electronic Systems, 2023" proposes a two-sided game strategy based on reinforcement learning against frequency-agile radars using masking signals. By optimizing time-domain resources, this algorithm can increase the jamming efficiency against frequency-agile radars. However, this model considers a relatively singular dimension, only considering the scheduling of the jammer's time-domain resources, and does not consider the situation where multiple networked frequency-agile radars exist in the environment. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides a reinforcement learning-based intelligent jamming method for networked frequency-agile radar systems, further enhancing the jamming capability against frequency-agile radars in electronic warfare environments.

[0007] The technical solution adopted in this invention is: a reinforcement learning intelligent jamming method for networked frequency-agile radar systems, the specific steps of which are as follows:

[0008] S1. Establish a scenario of multi-functional combined jammer countering networked frequency-agile radar, and determine the counter-relationship between the jammer and the frequency-agile radar.

[0009] S2. Design the jammer's operating mode, combining frequency sweeping suppression jamming with responsive suppression jamming, and sort out the schedulable resources of the jamming system.

[0010] S3. Analyze the radar time-varying parameters and determine that the radar operating mode is pulse group agile mode;

[0011] S4. Determine the evaluation indicators, that is, estimate the radar transmit signal-to-jamming ratio through the interceptor;

[0012] S5. Based on the evaluation indicators in step S4, establish a reward mechanism and transform the scheduling problem into an MDP process. In each action of the MDP process, select interception-interference time allocation and frequency sweep-response interference time allocation.

[0013] S6. Apply the deep deterministic strategy gradient algorithm to solve the MDP process and obtain the optimal scheduling strategy.

[0014] Furthermore, in step S1, the established scenario includes a networked frequency-agile radar system and a multi-functional jamming system.

[0015] In the radar system, the radars work together and transmit pulses at randomly varying frequencies; the multi-functional jamming system includes a programmable jamming signal transmitter and a signal interception receiver, both installed in the same spatial location.

[0016] Furthermore, step S2 is specifically as follows:

[0017] The interception and jamming systems operate in a time-division manner. For jamming, a pulse-by-pulse approach, also known as reactive jamming, is used, which involves intercepting and analyzing information from each radar pulse before sending a jamming pulse.

[0018] Based on pulse-by-pulse jamming, the jamming system will perform an "interception-jamming" task in each time interval. That is, when transmitting the nth (n = 1, 2, ...) pulse in the nth time interval, the jamming system will allocate interception time. and interference time Subject to the following constraints:

[0019]

[0020] Where T represents the pulse repetition interval.

[0021] Throughout the interception time Within this range, the jamming system monitors the entire operating frequency range of the remote aircraft, i.e., [f 1 ,f L Furthermore, the operating frequency range of the interference pulse also covers [f] 1 ,f L ].

[0022] Among them, f 1 with f L These represent the highest and lowest frequencies at which the radar operates, respectively, and L represents the number of available frequency points for the radar.

[0023] Meanwhile, the interference system will be in the memory buffer. The first N in the middle of the storage t The frequency band where radar signals were previously detected during the time interval That is, in During this period, the jamming system collects the frequency bands from which the interceptor detects radar signals into a set. In Chinese, the expression is as follows:

[0024]

[0025] Where, N t This indicates the retention time of data in the memory buffer. The interference system will utilize the memory buffer. The information in the data is used to implement interference strategies.

[0026] Assume there are I independent frequency-agile radars, denoted by R. i (i = 1, 2, ..., I) indicates that for each radar R i The pulse transmitted in the nth time interval is represented as

[0027] For a pulse with a transmission time interval of [(n-1)T, nT] Interference pulses will occur at time intervals Internal emission. The jamming pulse is based on two modes: scanning jamming and point-to-point jamming. The total power allocated to these two modes should not exceed the power limit of the jamming system, as expressed below:

[0028]

[0029] in, and P represents the power allocated to scanning interference and point-to-point interference, respectively. J This indicates the maximum power of the interference system.

[0030] (1) Scanning interference mode:

[0031] From During the interference time interval up to nT, the scanning interference will be in the frequency range [f] 1 ,f L The system uniformly emits interference pulses, which affect radar R. i Interference n s,i The expression for (t) is as follows:

[0032]

[0033] Among them, G t,j The gain of the interference pulse is represented by rect(x), which is a rectangular function; it takes the value 1 if x belongs to [0,1], and 0 otherwise. i Indicates radar R i Distance from the interfering system.

[0034] (2) Point-to-point interference mode:

[0035] Point-to-point jamming will continuously transmit jamming pulses, the carrier frequency of which exists in the memory buffer. In the middle, its radar R i Interference n d,i The expression for (t) is as follows:

[0036]

[0037] Here, bool(x) represents a Boolean function that equals 1 when the expression x is true, and 0 otherwise; f i (n) Indicates pulse agile frequency.

[0038] According to equations (4) and (5), the time-varying interference noise power n in the nth time interval is... J,i The expression for (t) is as follows:

[0039] n J,i (t)=P n0 +n s,i (t)+n d,i (t), (6)

[0040] in, This represents the power of ambient noise.

[0041] Furthermore, step S3 is specifically as follows:

[0042] The I independent frequency-agile radars all operate independently, periodically emitting pulses. These frequency-agile radars are all of the same type, with similar pulse repetition intervals, and all radars use the same frequency agility mode, namely pulse group agility mode.

[0043] For each radar R i The pulse transmitted in the nth time interval is represented as The expression for the kth pulse group of the radar is as follows:

[0044]

[0045] Among them, H i Indicates the total number of pulse transmissions. yes Another way to represent it is as abbreviated as pulse The agile frequency is expressed as f i (n,k) The expression is as follows:

[0046]

[0047] stfl =f l-1 +Δf(1≤l≤L), (9)

[0048] Where Δf represents the pulse frequency interval, which is a fixed frequency value. According to equations (7) and (8), the pulse... The time-domain expression is as follows:

[0049]

[0050] Where t represents the timestamp, φ(t) represents the pulse modulation function, τ represents the pulse width, and P i Indicates radar R i The transmission power; This indicates rounding down for x.

[0051] Furthermore, step S4 is specifically as follows:

[0052] The evaluation metric is determined, and the new interference-to-noise ratio (ESINR) of the transmitted signal is defined as the ratio of radar pulse power to interference power at the transmission point. The ESINR expression for the radar in the nth time interval is as follows:

[0053]

[0054] Where k0 represents the Boltzmann constant, and T0 represents the ambient temperature. t F represents the antenna gain of the radar. r This represents the beamforming direction factor. J,i (t) represents the time-varying interference noise power involved in step S2. For radar R i When ESINR When the value drops below the threshold φ, it is considered to be radar R. i Successfully interfered with.

[0055] For each pulse i = 1, 2, ..., I, and its transmission time period is [(n-1)T, nT]. The intercepting receiver will [within the time period]... Internal monitoring frequency range [f 1 ,f L At the intercept receiver, the pulse The cumulative signal-to-noise ratio (SNR) is calculated as follows:

[0056]

[0057] Among them, G j This indicates the intercepted receiver gain; This represents the noise power at the interceptor. Based on equation (12), the noise power at the interceptor for pulses is estimated. Detection probability The expression is as follows:

[0058]

[0059] Where, p fa This represents the false alarm probability of intercepting the receiver. During this period, the jamming system collects the frequency bands from which the interceptor detects radar signals into a set. In the process of intercepting a radar pulse, the jamming system estimates radar parameters from the intercepted signal, including position, direction of arrival, and transmit power. Based on these radar parameters, the jamming system approximates the radar's ESINR and evaluates the jamming effect accordingly.

[0060] Furthermore, step S5 is specifically as follows:

[0061] The attack of a frequency-agile radar by a jamming system is modeled as a Markov decision process (MDP), which includes four elements: state S, action A, reward R, and policy P.

[0062] State S:

[0063] During the nth pulse transmission, the interceptor detects the operating frequency range and calculates the memory buffer. The state of the interference system at the nth operation is defined as follows:

[0064]

[0065] in, This indicates the current number of frequency points in the memory buffer, while This represents the number of radars estimated by the jamming system before the nth operation.

[0066] Action A:

[0067] During the nth operation, the interference system allocates time and power resources to allocate interference resources. The expression for action A is as follows:

[0068]

[0069] The total operating time of the jamming system consists of the interception time. and interference time Coverage, among which The power resources of the jammer are determined by the scanning jamming power. Point-to-point interference power Consumption, i.e.

[0070] Reward R:

[0071] The jamming system calculates R for each radar in the nth time interval.i Transmit signal-to-interference ratio Where i∈{1,2,...,I}. The reward function of the jamming system is defined as the reduction ratio of the transmit signal-to-interference-plus-noise ratio (SNR) of all intercepted frequency-agile radars, and the expression is as follows:

[0072]

[0073] The interceptor receiver locates the intercepted radar based on the received radar pulses and estimates the total number of radars in the environment based on the location information. When the interceptor receiver first detects any radar in the surrounding area, the jamming system updates the estimated number of radars. The transmit signal-to-interference-plus-noise ratio (SINNR) of the radar when it is unaffected by jamming systems is expressed as follows:

[0074]

[0075] Based on this reward function, the jamming system will minimize the performance of all detected radars in the environment, with the ultimate goal of achieving the most effective jamming strategy.

[0076] Furthermore, step S6 is specifically as follows:

[0077] The Markov decision model established in step S5 has a continuous action space. The deep deterministic policy gradient algorithm is used to solve this Markov decision model, and the policy function μ... θ (s) is used to represent the policy, and the state s (n) Mapping to deterministic action a (n) The expression is as follows:

[0078] a (n) =μ θ (s (n) (18)

[0079] Wherein, the policy function μ θ An approximation is made using a neural network, which is called an actor network, and the network parameters of the actor network are θ. μ .

[0080] Value function Q μ (s (n) ,a (n) The value denoted by is the expected reward obtained when executing policy μ. This is approximated by another neural network function, Q. μ (s (n) ,a (n) This neural network is called the Critics Network Q. θ (s (n) ,a (n) ), and the network parameters of the critic network are θ. QWhen updating the commentator network, the loss function expression is as follows:

[0081]

[0082] Where r(s) (n) ,a (n) ) indicates that by s (n) and a (n) The calculated reward value, where γ represents the discount factor. The loss function expression for the actor network is as follows:

[0083]

[0084] That is, in each iteration of training, the Adam optimization algorithm is used to optimize the parameters θ in the two neural networks respectively. Q and θ μ The optimal policy a corresponding to the Markov decision model can then be solved. (n) =μ θ (s (n) ).

[0085] The beneficial effects of this invention are as follows: First, the method of this invention establishes a multi-radar versus jammer confrontation scenario and designs a pulse-by-pulse jamming system model that combines point jamming and frequency sweeping jamming. Then, the time-domain and energy-domain resource allocation decision-making process of the jamming system when jamming a radar group is modeled as a Markov decision process, and a deep deterministic policy gradient algorithm is used to train and solve for the optimal policy of the jamming system. The optimized policy obtained by the method of this invention has strong adaptability, good model scalability, and high resource utilization efficiency. It can estimate the operating mode and number of radar systems to be jammed, and perform adaptive jamming based on this. By adaptively adjusting the operating parameters of the jammer's time-domain and frequency-domain resources, compared with existing jamming methods that rely on specific rules for frequency-agile radars, it can significantly enhance the jamming efficiency against networked frequency-agile radar systems with various configurations, and improve the jammer's mission performance. Simultaneously, deep reinforcement learning algorithms are used to optimize the combined suppression jamming strategy, enhancing the jammer's strike capability in electronic warfare environments. Attached Figure Description

[0086] Figure 1 This is a flowchart of a reinforcement learning intelligent jamming method for networked frequency-agile radar systems according to the present invention.

[0087] Figure 2 This is a schematic diagram of a counter-environment scenario between an interference system and a frequency-agile radar group in an embodiment of the present invention.

[0088] Figure 3 This is a schematic diagram illustrating the interaction between the jamming system and the radar environment in an embodiment of the present invention.

[0089] Figure 4 This diagram illustrates the working details of the interference system in an embodiment of the present invention.

[0090] Figure 5 This is a comparison chart of the total number of successful jamming attempts under different numbers of radars in this embodiment of the invention.

[0091] Figure 6 This is a comparison chart of various strategies for the average transmit signal-to-interference-plus-noise ratio of radars under different numbers of radars in this embodiment of the invention. Detailed Implementation

[0092] This invention is primarily verified using simulation experiments, and all steps and conclusions have been validated correctly using Python 3.7. The method of this invention will be further explained below with reference to the accompanying drawings and embodiments.

[0093] To facilitate the description of the method of the present invention, the following terms will first be explained:

[0094] Term 1: Frequency-agile radar;

[0095] Frequency-agile radar is a type of radar system that operates based on frequency modulation and frequency scanning. In frequency-agile radar, the signal transmitted by the radar periodically changes its frequency; this frequency change can be continuous, abrupt, or random.

[0096] Term 2: Suppression of interference;

[0097] Suppression jamming is an electronic warfare tactic designed to interfere with, confuse, or mask the receiver of a radar system by transmitting jamming signals, preventing it from correctly detecting targets or extracting target information.

[0098] Term 3: Intercepting the receiver;

[0099] An intercept receiver is a passive electronic countermeasures device used to intercept and monitor signals transmitted by radar systems. It works by receiving signals emitted by a target radar system, demodulating and analyzing them to obtain relevant information about the target radar system, such as frequency, pulse repetition frequency, and modulation method.

[0100] like Figure 1 The flowchart shown below illustrates a reinforcement learning-based intelligent jamming method for networked frequency-agile radar systems according to the present invention. The specific steps are as follows:

[0101] S1. Establish a scenario of multi-functional combined jammer countering networked frequency-agile radar, and determine the counter-relationship between the jammer and the frequency-agile radar.

[0102] S2. Design the jammer's operating mode, combining frequency sweeping suppression jamming with responsive suppression jamming, and sort out the schedulable resources of the jamming system.

[0103] S3. Analyze the radar time-varying parameters and determine that the radar operating mode is pulse group agile mode;

[0104] S4. Determine the evaluation indicators, that is, estimate the radar transmit signal-to-jamming ratio through the interceptor;

[0105] S5. Based on the evaluation indicators in step S4, establish a reward mechanism and transform the scheduling problem into an MDP process. In each action of the MDP process, select interception-interference time allocation and frequency sweep-response interference time allocation.

[0106] S6. Apply the Deep Deterministic Policy Gradient Algorithm (DDPG) to solve the MDP process and obtain the optimal scheduling policy.

[0107] In this embodiment, in step S1, the established scenario includes a networked frequency-agile radar system and a multi-functional jamming system.

[0108] In the radar system, the radars work together and transmit pulses at randomly varying frequencies; the multi-functional jamming system includes a programmable jamming signal transmitter and a signal interception receiver, both installed in the same spatial location.

[0109] In this embodiment, step S2 is specifically as follows:

[0110] When a programmable jamming transmitter (jammer) emits jamming signals, it is not feasible to completely isolate the interference to the signal interception receiver, making it impossible to perform jamming and interception tasks simultaneously. Therefore, the interception and jamming system must operate in a time-division manner.

[0111] For the jamming function, this embodiment employs a pulse-by-pulse approach, also known as reactive jamming. That is, before sending a jamming pulse, the information of each radar pulse is intercepted and analyzed. Based on the pulse-by-pulse jamming mode, the jamming system will perform an "interception-jamming" task in each time interval. Specifically, when sending the nth (n = 1, 2, ...) pulse in the nth time interval, the jamming system will allocate interception time. and interference time Subject to the following constraints:

[0112]

[0113] Where T represents the pulse repetition interval.

[0114] Throughout the interception time Within this range, the jamming system monitors the entire operating frequency range of the remote aircraft, i.e., [f 1 ,f L Furthermore, the operating frequency range of the interference pulse also covers [f]1 ,f L ].

[0115] Among them, f 1 with f L These represent the highest and lowest frequencies at which the radar operates, respectively, and L represents the number of available frequency points for the radar.

[0116] Meanwhile, the interference system will be in the memory buffer. The first N in the middle of the storage t The frequency band where radar signals were previously detected during the time interval That is, in During this period, the jamming system collects the frequency bands from which the interceptor detects radar signals into a set. In Chinese, the expression is as follows:

[0117]

[0118] Where, N t This indicates the retention time of data in the memory buffer. The interference system will utilize the memory buffer. The information in the data is used to implement interference strategies.

[0119] Assume there are I independent frequency-agile radars, denoted by R. i (i = 1, 2, ..., I) indicates that for each radar R i The pulse transmitted in the nth time interval is represented as

[0120] For a pulse with a transmission time interval of [(n-1)T, nT] Interference pulses will occur at time intervals Internal emission. The jamming pulse is based on two modes: scanning jamming and point-to-point jamming. The total power allocated to these two modes should not exceed the power limit of the jamming system, as expressed below:

[0121]

[0122] in, and P represents the power allocated to scanning interference and point-to-point interference, respectively. J This indicates the maximum power of the interference system.

[0123] (1) Scanning interference mode:

[0124] From During the interference time interval up to nT, the scanning interference will be in the frequency range [f] 1 ,f L The system uniformly emits interference pulses, which affect radar R. i Interference n s,iThe expression for (t) is as follows:

[0125]

[0126] Among them, G t,j The gain of the interference pulse is represented by rect(x), which is a rectangular function; it takes the value 1 if x belongs to [0,1], and 0 otherwise. i Indicates radar R i Distance from the interfering system.

[0127] (2) Point-to-point interference mode:

[0128] Point-to-point jamming will continuously transmit jamming pulses, the carrier frequency of which exists in the memory buffer. In the middle, its radar R i Interference n d,i The expression for (t) is as follows:

[0129]

[0130] Here, bool(x) represents a Boolean function that equals 1 when the expression x is true, and 0 otherwise; f i (n) Indicates pulse agile frequency.

[0131] According to equations (4) and (5), the time-varying interference noise power n in the nth time interval is... J,i The expression for (t) is as follows:

[0132]

[0133] in, This represents the power of ambient noise.

[0134] In this embodiment, step S3 is specifically as follows:

[0135] Each of the aforementioned independent frequency-agile radars operates independently, periodically emitting pulses. These frequency-agile radars are all of the same type, with similar pulse repetition intervals. All radars employ the same frequency agility mode, namely pulse group agility mode. This mode maintains a constant carrier frequency within a set of pulse trains, while different groups can occupy different frequencies. Within a pulse group, pulses are processed consistently to achieve better mission performance.

[0136] For each radar R i The pulse transmitted in the nth time interval is represented as The expression for the kth pulse group of the radar is as follows:

[0137]

[0138] Among them, H i This represents the total number of pulse transmissions (from the nth to the (n+Hth)th pulse). i -1 pulse), yes Another representation, to emphasize its association with pulse group k, is abbreviated as pulse The agile frequency is expressed as f i (n,k) The expression is as follows:

[0139]

[0140] stf l =f l-1 +Δf(1≤l≤L), (9)

[0141] Where Δf represents the pulse frequency interval, which is a fixed frequency value. According to equations (7) and (8), the pulse... The time-domain expression is as follows:

[0142]

[0143] Where t represents the timestamp, φ(t) represents the pulse modulation function, τ represents the pulse width, and P i Indicates radar R i The transmission power; This indicates rounding down for x.

[0144] In this embodiment, step S4 is specifically as follows:

[0145] To conveniently estimate radar performance and determine evaluation metrics from the perspective of the jamming system, a new interference-to-noise ratio (ESINR) is defined as the ratio of radar pulse power to jamming power at the transmission point. The ESINR expression for the radar in the nth time interval is as follows:

[0146]

[0147] Where k0 represents the Boltzmann constant, and T0 represents the ambient temperature. t F represents the antenna gain of the radar. r This represents the beamforming direction factor. J,i (t) represents the time-varying interference noise power involved in step S2. For radar R i When ESINR When the value drops below the threshold φ (φ can be set according to the actual interference effect requirements; the higher the actual interference effect requirements, the smaller the φ should be), it is considered to be radar R. i Successfully interfered with.

[0148] For each pulse i = 1, 2, ..., I, and its transmission time period is [(n-1)T, nT]. The intercepting receiver will [within the time period]... Internal monitoring frequency range [f 1 ,f L At the intercept receiver, the pulse The cumulative signal-to-noise ratio (SNR) is calculated as follows:

[0149]

[0150] Among them, G j This indicates the intercepted receiver gain; This represents the noise power at the interceptor. Based on equation (12), the noise power at the interceptor for pulses is estimated. Detection probability The expression is as follows:

[0151]

[0152] Where, p fa This represents the false alarm probability of intercepting the receiver. During this period, the jamming system collects the frequency bands from which the interceptor detects radar signals into a set. In the process of intercepting a radar pulse, the jamming system estimates radar parameters from the intercepted signal, including position, direction of arrival, and transmit power. Based on these radar parameters, the jamming system approximates the radar's ESINR and evaluates the jamming effect accordingly.

[0153] In this embodiment, step S5 is specifically as follows:

[0154] The attack of a frequency-agile radar by a jamming system is modeled as a Markov decision process (MDP), which is a sequential decision process with Markov properties, including four elements: state S, action A, reward R, and policy P.

[0155] State S:

[0156] During the nth pulse transmission, the interceptor detects the operating frequency range and calculates the memory buffer. The state of the interference system at the nth operation is defined as follows:

[0157]

[0158] in, This indicates the current number of frequency points in the memory buffer, while This represents the number of radars estimated by the jamming system before the nth operation.

[0159] Action A:

[0160] During the nth operation, the interference system allocates time resources (i.e., interception time). and interference time ) and power resources (i.e., scanning interference power) Point-to-point interference power To allocate interfering resources. The expression for action A is as follows:

[0161]

[0162] The total operating time of the jamming system consists of the interception time. and interference time Coverage, among which The power resources of the jammer are determined by the scanning jamming power. Point-to-point interference power Consumption, i.e.

[0163] Reward R:

[0164] The jamming system calculates R for each radar in the nth time interval. i Transmit signal-to-interference ratio Where i∈{1,2,...,I}. The reward function of the jamming system is defined as the reduction ratio of the transmit signal-to-interference-plus-noise ratio (SNR) of all intercepted frequency-agile radars, and the expression is as follows:

[0165]

[0166] The interceptor receiver locates the intercepted radar based on received radar pulses and estimates the total number of radars in the environment based on the location information. When the interceptor receiver first detects any radar in the surrounding area, the jamming system updates the estimated number of radars. The transmit signal-to-interference-plus-noise ratio (SINNR) of the radar when it is unaffected by jamming systems is expressed as follows:

[0167]

[0168] Based on this reward function, the jamming system will minimize the performance of all detected radars in the environment, with the ultimate goal of achieving the most effective jamming strategy.

[0169] In this embodiment, step S6 is specifically as follows:

[0170] The Markov decision model established in step S5 has a continuous action space. The deep deterministic policy gradient algorithm is used to solve this Markov decision model, and the policy function μ... θ (s) is used to represent the policy, and the state s (n) Mapping to deterministic action a(n) The expression is as follows:

[0171] a (n) =μ θ (s (n) (18)

[0172] Wherein, the policy function μ θ An approximation is made using a neural network, which is called an actor network, and the network parameters of the actor network are θ. μ .

[0173] Value function Q μ (s (n) ,a (n) The value denoted by is the expected reward obtained when executing policy μ. This is approximated by another neural network function, Q. μ (s (n) ,a (n) This neural network is called the Critics Network Q. θ (s (n) ,a (n) ), and the network parameters of the critic network are θ. Q When updating the commentator network, the loss function expression is as follows:

[0174]

[0175] Where r(s) (n) ,a (n) ) indicates that by s (n) and a (n) The calculated reward value, where γ represents the discount factor. The loss function expression for the actor network is as follows:

[0176]

[0177] That is, in each iteration of training, the Adam optimization algorithm is used to optimize the parameters θ in the two neural networks respectively. Q and θ μ The optimal policy a corresponding to the Markov decision model can then be solved. (n) =μ θ (s (n) ).

[0178] As the actor network and commentator network are updated, after sufficient exploration of the action space, the ε-greedy policies corresponding to the actor network and commentator network tend to converge, yielding the optimal action policy for this decision-making process, which is also the optimal solution to the scheduling problem.

[0179] In this embodiment, the environmental parameters, deep deterministic policy gradient algorithm parameters, and interference system parameters are shown in Table 1.

[0180] Table 1

[0181]

[0182] like Figure 2 As shown in the diagram, this embodiment of the invention illustrates a scenario where a jamming system and a frequency-agile radar group are engaged in combat. Multiple radars are positioned against a single jammer, while the jammer moves in the airspace. Each radar performs frequency agility according to the pulse group agility principle, and the operating parameters of each radar are different. The jamming system performs jamming in a responsive manner, intercepting each radar's pulse transmission interval before jamming. The jamming methods include point jamming and frequency sweeping jamming.

[0183] Figure 3 This is a schematic diagram illustrating the interaction between the jamming system and the radar environment in an embodiment of the present invention. Figure 4 This is a diagram illustrating the working details of the interference system in an embodiment of the present invention, wherein... Figure 4 (a) is a detailed diagram of the frequency domain operation of the interference system. Figure 4 (b) is a diagram showing the details of the resource allocation of the jamming system. The jamming system will counter a networked radar system consisting of four frequency-agile radars in the environment. As shown in the diagram, the jamming system will adjust the allocation of time-domain and energy-domain resources based on the intercepted information in order to reduce the radar's transmit signal-to-interference-plus-noise ratio as much as possible.

[0184] Figure 5 This is a comparison chart of the total number of successful jamming attempts under different numbers of radars in the embodiments of the present invention. The jamming success rate of the strategy designed by the method of the present invention is higher than that of existing fixed strategies and random strategies in environments with various numbers of radars. Figure 6 This is a comparison chart of the average transmit signal-to-interference-plus-noise ratio (SINORR) of radars under different numbers of radars in embodiments of the present invention. The method of the present invention designs a strategy that results in a higher radar performance degradation rate than existing fixed and random strategies in environments with various numbers of radars.

[0185] In summary, the method of this invention can effectively optimize the intelligent jamming strategy of jamming systems against networked frequency-agile radars, enhancing the jamming performance against frequency-agile radars and improving the jamming efficiency of jamming systems in electronic warfare environments. This invention overcomes the problems of low jamming success rate and inflexible resource scheduling of existing jamming strategies against networked frequency-agile radar systems. The optimized strategy obtained by the method of this invention has strong adaptability and high resource utilization efficiency. Compared with existing jamming methods that rely on specific rules against frequency-agile radars, it can significantly enhance the jamming efficiency against networked frequency-agile radar systems with various configurations, explore the mission performance of jammers, and simultaneously utilize deep reinforcement learning algorithms to optimize combined suppression jamming strategies, thereby improving the jammer's strike capability in electronic warfare environments.

[0186] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. A reinforcement learning-based intelligent jamming method for networked frequency-agile radar systems, comprising the following steps: S1. Establish a scenario of multi-functional combined jammer countering networked frequency-agile radar, and determine the counter-relationship between the jammer and the frequency-agile radar. S2. Design the jammer's operating mode, combining frequency sweeping suppression jamming with responsive suppression jamming, and sort out the schedulable resources of the jamming system. S3. Analyze the radar time-varying parameters and determine that the radar operating mode is pulse group agile mode; S4. Determine the evaluation indicators, that is, estimate the radar transmit signal-to-jamming ratio through the interceptor; S5. Based on the evaluation indicators in step S4, establish a reward mechanism and transform the scheduling problem into an MDP process. In each action of the MDP process, select interception-interference time allocation and frequency sweep-response interference time allocation. S6. Apply the Deep Deterministic Strategy Gradient Algorithm to solve the MDP process and obtain the optimal scheduling strategy; In step S1, the established scenario includes a networked frequency-agile radar system and a multi-functional jamming system. in, The radar system operates in coordination with the radars and transmits pulses at randomly varying frequencies; the multi-functional jamming system includes a programmable jamming signal transmitter and a signal interception receiver, both installed in the same spatial location. Step S2 is as follows: The interception and jamming system operates in a time-division manner; for jamming, it uses a pulse-by-pulse approach, also known as reactive jamming, which means intercepting and analyzing the information of each radar pulse before sending jamming pulses. Based on pulse-by-pulse jamming mode, the jamming system will perform an "interception-jamming" task in each time interval, that is, in the pulse-by-pulse jamming mode... Send the first time interval During the pulse, the jamming system will allocate interception time. and interference time Subject to the following constraints: (1); in, Indicates the pulse repetition interval; ; Throughout the interception time Within this range, the jamming system monitors the entire operating frequency range of the remote aircraft, i.e. Furthermore, the operating frequency range of the interference pulse also covers ; in, and These represent the highest and lowest frequencies at which the radar operates. Indicates the number of available frequency points for the radar; Meanwhile, the interference system will be in the memory buffer. Before storage The frequency band where radar signals were previously detected during the time interval That is, in During this period, the jamming system collects the frequency bands from which the interceptor detects radar signals into a set. In Chinese, the expression is as follows: (2); in, This indicates the retention time of data in the memory buffer; the interference system will utilize the memory buffer. Use the information in the data to implement interference strategies; There are a total of settings Taiwan's independent frequency-agile radar, using express, For each radar In the The pulse transmitted over a time interval is represented as ; For the transmission time interval is pulse The interference pulse will be at the time interval Internal emission; the jamming pulse is based on two modes: scanning jamming and point-to-point jamming; the total power allocated to these two modes should not exceed the power limit of the jamming system, as expressed below: (3); in, and These represent the power allocated to scanning interference and point-to-point interference, respectively. Indicates the maximum power of the interfering system; (1) Scanning interference mode: From arrive During the interference time interval, the scanning interference will be in the frequency range Uniformly emitted interference pulses within the radar, which affect the radar Interference The expression is as follows: (4); in, Indicates the gain of the interference pulse. Represents a rectangular function, if belong If the result is positive, then take 1; otherwise, take 0. Indicates radar Distance to the jamming system; (2) Point-to-point interference mode: Point-to-point jamming will continuously transmit jamming pulses, the carrier frequency of which exists in the memory buffer. In the middle, its radar Interference The expression is as follows: (5); in, Represent a Boolean function, where the expression ... It equals 1 if true, otherwise it equals 0; Indicates pulse agile frequency; According to equations (4) and (5), the first Time-varying interference noise power of time interval The expression is as follows: (6); in, This represents the power of ambient noise.

2. The reinforcement learning intelligent jamming method for networked frequency-agile radar systems according to claim 1, characterized in that, Step S3 is as follows: The Each independent frequency-agile radar operates independently, periodically emitting pulses. These frequency-agile radars are all of the same type, with similar pulse repetition intervals, and all radars use the same frequency agility mode, namely pulse group agility mode. For each radar In the The pulse transmitted over a time interval is represented as Then the radar's first The expression for the pulse group is as follows: (7); in, Indicates the total number of pulse transmissions. yes Another way to represent it is as abbreviated as ;pulse The agile frequency is expressed as The expression is as follows: (8); (9); in, This represents the pulse frequency interval, which is a fixed frequency value; according to equations (7) and (8), the pulse... The time-domain expression is as follows: (10); in, Represents a timestamp. Represents the pulse modulation function. Indicates the pulse width. Indicates radar The transmission power; Indicates for Round down to the nearest integer.

3. The reinforcement learning intelligent jamming method for networked frequency-agile radar systems according to claim 2, characterized in that, Step S4 is as follows: The evaluation index is determined, and the new interference-to-noise ratio (ESINR) of the transmitted signal is defined as the ratio of the radar pulse power to the interference power at the transmission point; the radar at the [missing information]... The ESINR expression for the time interval is as follows: (11); in, Represents the Boltzmann constant. Indicates ambient temperature; This indicates the antenna gain of the radar. Indicates the beamforming direction factor; This indicates the time-varying interference noise power involved in step S2; for radar When ESINR Drop to threshold The following are considered radar. Successfully interfered with; For each pulse Its transmission time period is The intercepted receiver will be within a certain time period. Internal monitoring frequency range At the intercept receiver, the pulse The cumulative signal-to-noise ratio (SNR) is calculated as follows: (12); in, This indicates the intercepted receiver gain; Represents the noise power at the interceptor; based on equation (12), estimate the noise power at the interceptor for pulses. Detection probability The expression is as follows: (13); in, This indicates the false alarm probability of intercepting the receiver; in During this period, the jamming system collects the frequency bands from which the interceptor detects radar signals into a set. In the middle; after the pulse is intercepted, the jamming system estimates radar parameters from the intercepted signal, including: position, direction of arrival and transmission power; based on the radar parameters, the jamming system approximates the radar's ESINR and evaluates the jamming effect accordingly.

4. The reinforcement learning intelligent jamming method for networked frequency-agile radar systems according to claim 3, characterized in that, Step S5 is as follows: The attack of a frequency-agile radar by a jamming system is modeled as a Markov decision process (MDP), which includes four elements: state ,action ,award and strategy ; state : In the During the transmission of the next pulse, the interceptor detects the operating frequency range and calculates the memory buffer. ;No. The state of the interference system at the time of the next operation is defined as follows: (14); in, This indicates the current number of frequency points in the memory buffer, while Indicates the first The number of radars estimated by the jamming system before the next operation; action : In the During this operation, the interference system allocates time and power resources to allocate interference resources; actions The expression is as follows: (15); The total operating time of the jamming system consists of the interception time. and interference time Coverage, among which The power resources of the jammer are determined by the scanning jamming power. Point-to-point interference power Consumption, i.e. ; award : The interference system in the first Time interval calculation for each radar Transmit signal-to-interference ratio ,in The reward function for the jamming system is defined as the percentage reduction in the transmit signal-to-interference-plus-noise ratio (SNR) of all intercepted frequency-agile radars, expressed as follows: (16); The interceptor receiver locates the intercepted radar based on the received radar pulses and estimates the total number of radars in the environment based on the location information. When the interceptor receiver first detects any radar in the surrounding area, the jamming system updates the estimated number of radars. ; The transmit signal-to-interference-plus-noise ratio (SINNR) of the radar when it is unaffected by jamming systems is expressed as follows: (17); Based on this reward function, the jamming system will minimize the performance of all detected radars in the environment, with the ultimate goal of achieving the most effective jamming strategy.

5. The reinforcement learning intelligent jamming method for networked frequency-agile radar systems according to claim 4, characterized in that, Step S6 is as follows: The Markov decision model established in step S5 has a continuous action space. The deep deterministic policy gradient algorithm is used to solve this Markov decision model, and the policy function... Used to represent the strategy, the state Mapping to deterministic actions The expression is as follows: (18); Among them, the policy function An approximation is made using a neural network, which is called an actor network, and the network parameters of the actor network are... ; Value function Indicates the execution strategy The expected reward obtained at that time; approximated by another neural network function. This neural network is called the critic network. And the network parameters of the critic network are When updating the commentator network, the loss function expression is as follows: (19); in, Indicates by and The calculated reward value; and Let represent the discount factor; the loss function expression for the actor network is as follows: (20); That is, in each iteration of training, the Adam optimization algorithm is used to optimize the parameters in the two neural networks respectively. and The optimal strategy corresponding to the Markov decision model can then be solved. .

Citation Information

Patent Citations

  • Radar interference game strategy design method based on neural network virtual self-game

    CN114236477A

  • Radar space-time-frequency-energy multi-domain combined intelligent active anti-interference method

    CN115932750A