Traction interference time domain strategy generation method based on fuzzy actor commentator algorithm
Through the target echo model and interference signal model established by the fuzzy actor critic algorithm, the fuzzy inference system is used to accurately process the radar PRI, which solves the problem of interference strategy generation under the condition of incomplete radar information, realizes active traction of radar wave gates and time-domain spoofing interference, and improves the intelligence level of radar confrontation.
Patent Information
- Application Number
- CN202510550974.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-29
AI Technical Summary
Under the condition that existing radar jamming technology obtains radar information in incomplete conditions, it is difficult to generate effective interference strategies in high-dimensional continuous state space, resulting in poor confrontation effect.
The target echo model and interference signal model are established based on the fuzzy actor critic algorithm, and the fuzzy inference system is used to accurately process the radar PRI incomplete information, establish a generalized Markov decision-making process, and solve the traction interference time domain strategy through the fuzzy actor critic algorithm to realize active traction of the radar wave gate.
Under the condition of incomplete radar information, the radar wave gate is effectively pulled away from the real target echo signal, achieving the time-domain spoofing interference effect, and improving the intelligence level of radar confrontation.
Smart Images

Figure CN120385979A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of radar jamming, and particularly relates to a time-domain strategy generation method for towing jamming based on a fuzzy actor-critic algorithm. Background Art
[0002] Tracking radars have a wide range of applications and can achieve target recognition, measurement of target motion information, and tracking of target trajectory information, which can pose a serious threat to the survival of high-value targets. As a common electronic countermeasure technology, the range gate pull-off jamming method deceives the radar's detected target range by dragging the radar range gate away from the true target echo signal. With the development of radar signal processing, electronic information technology, and artificial intelligence technology, traditional jamming means based on fixed templates have gradually become ineffective and are difficult to cope with the dynamic time-varying characteristics of electronic countermeasures. Therefore, researching a new and effective jamming technology is an urgent need for future electronic countermeasure applications.
[0003] With the continuous improvement of the intelligence and informatization levels of radars and the development of cognitive intelligent radars, many radar systems already have the ability of autonomous learning, and the combat effectiveness of existing electronic jamming methods against radars is declining. Therefore, it is necessary to consider the intelligence of jamming methods and realize the transformation from "pre-programmed response" to "cognitive countermeasure". The literature "Feng L W, Liu S T, Xu H Z. Multifunctional radar cognitive jamming decision based on dueling double deep Q-network. IEEE Access, 2022, 10: 112150 - 112157" proposed an interference decision optimization algorithm based on a double deep Q-network for multifunctional radars. This algorithm uses a value function reflecting the change of radar state and an advantage function related to radar state and interference actions to improve the cognitive jamming level against unknown radar modes, and then uses a dueling network for interference strategy selection and effect evaluation to further improve the accuracy of decision-making. The literature "Zhang Baikai, Zhu Weigang. DQN cognitive jamming decision method for multifunctional radars. Systems Engineering and Electronics, 2020, 42(04): 819 - 825" proposed a deep Q-network interference decision algorithm for multifunctional radars, which realizes intelligent interference decision-making by constructing an interference library and autonomously learning interference strategies, but the decision-making accuracy still needs to be improved. Although the jamming methods in the above literature have better effects compared with existing radar countermeasure methods. However, they only consider jamming decision-making under the condition that the jammer can obtain all or accurate information of the radar. Currently, there is little research on the method for generating interference strategies in a high-dimensional continuous state space under the condition of incomplete radar information acquisition. Therefore, it is necessary to conduct research on it. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a method for generating a time-domain strategy of towing interference based on a fuzzy actor-critic algorithm, which realizes the effect of time-domain deception interference and effectively towes the radar gate away from the real target echo signal.
[0005] The technical solution adopted by the present invention is as follows: a method for generating a time-domain strategy of towing interference based on a fuzzy actor-critic algorithm, and the specific steps are as follows:
[0006] S1. According to the confrontation scenario between the radar and the jammer, establish a target echo model and a jamming signal model;
[0007] S2. Based on the confrontation mechanism between the towing interference and the radar, establish a mathematical optimization model for generating the towing interference strategy;
[0008] S3. Regard the jammer as an intelligent agent, use the fuzzy inference system to precisely process the incomplete information of the radar PRI perceived by the jammer, obtain the time when the radar pulse arrives at the receiver as the state, and establish a generalized Markov decision process for the confrontation between the towing interference and the radar;
[0009] S4. Use the fuzzy actor-critic algorithm to solve the established generalized Markov decision problem, obtain the strategy selection of the towing interference emission time, and ensure that the jammer actively towes the radar range gate under the condition of incomplete radar information acquisition.
[0010] Furthermore, the specific content of step S1 is as follows:
[0011] During the confrontation process, the radar emits a pulse train to irradiate the target. The number of pulses emitted within each coherent processing interval CPI is M, and it repeats continuously with the pulse repetition interval PRI. Then the radar emits the m-th pulse signal s c with the following expression: m (t)
[0012] s m (t) = A R s(t - mT P,R,m ) exp(j2πf c t) (1)
[0013] where m = 1, 2,..., M; t represents time, A R represents the amplitude of the radar emission signal, T P,R,m represents the PRI of the m-th pulse, and s(t) represents a linear frequency modulation signal LFM. The specific waveform expression is as follows:
[0014]
[0015] where T d,R represents the pulse width, and μ = B / T d,RLet \(k_{FM}\) denote the frequency modulation slope, \(B\) denote the signal bandwidth, and \(rect(\cdot)\) denote the rectangular function. The expression is as follows:
[0016]
[0017] The jammer can sense the moment when the radar pulse signal is intercepted by the jammer. When the jammer intercepts the \(m\)-th radar pulse signal \(x\) m (t), the expression is as follows:
[0018] x m (t)=\(\alpha\) m s m (t - t R ), (4)
[0019] where \(\alpha\) m denotes the propagation path loss, \(t\) R =\(R(t) / c\) represents the time delay generated during the propagation of the LFM signal, \(R(t)\) represents the distance function between the target and the radar, and \(c\) represents the speed of light.
[0020] Then, the jammer measures the parameters of the radar pulse signal. The parameter measurement includes: the center frequency of the radar signal, the signal bandwidth, direction finding and angle measurement of the radar position.
[0021] By sensing the radar pulse emission law, before the radar pulse reaches the jammer receiver, the jammer emits a range forward towing interference signal in advance and controls the forwarding delay to gradually decrease, simulating the target approaching the radar to exert pressure on the radar side. Then the jammer emits a range towing interference signal \(J\) m (t J,m ), and the expression is as follows:
[0022] J m (t J,m ) = \(A\) J \(\alpha\) m s m (t J,m - t R + \(\Delta t\) J (t)) (5)
[0023] where \(A\) J denotes the amplitude of the range towing interference signal emitted by the jammer, \(t\) J,m denotes the emission moment of the \(m\)-th pulse of the range forward towing interference, and \(A\) J > \(A\) R , \(\Delta t\) J (t) represents the forwarding delay function of the range towing interference.
[0024] Set \(v\) J to denote the speed during uniform towing. Then the forwarding delay function \(\Delta t\) of the range towing interferenceJ (t) The expression is as follows:
[0025]
[0026] Among them, T J represents the total time of a single traction interference. It is divided into three parts according to the three stages of the traction interference. t1 represents the cut-off moment of the interference capture period, and t2 represents the cut-off moment of the interference traction period.
[0027] Furthermore, the specific steps of step S2 are as follows:
[0028] The mathematical optimization model expression for generating the traction interference strategy is as follows:
[0029]
[0030] Among them, t J,m represents the emission moment of the m-th range traction interference pulse, represents the deduced m-th radar PRI, t min ≤t J,m ≤t max represents all interference pulses received within the defined time window of the radar. t min and t max represent the minimum and maximum values of the time window range respectively; represents the specific range of the jammer's deduced radar PRI. T min represents the minimum value of the jammer's deduced radar PRI, and T max represents the maximum value of the jammer's deduced radar PRI. represents the number of consecutive traction target echo CPIs for the range forward traction interference. The specific expression is as follows:
[0031]
[0032] Among them, represents the time-domain overlap probability between the range forward traction interference pulse and the radar pulse. τ represents the probability threshold. When it indicates that the range forward traction interference can accurately deduce the radar signal PRI information. Its specific expression is as follows:
[0033]
[0034] Among them, represents the time-domain range of the m-th radar pulse. t R,m-1 represents the moment when the jammer intercepts the (m - 1)-th radar pulse, represents the time-domain range of the m-th range forward traction interference pulse. T d,J represents the width of the interference pulse. Then t J,mThe specific expression is as follows:
[0035]
[0036] Further, the specific steps of step S3 are as follows:
[0037] S31. Use a fuzzy inference system to precisely process the incomplete information of the radar PRI sensed by the jammer;
[0038] Adopt a multi-input single-output (MISO) fuzzy inference system (FIS). Utilize the function approximation ability of the fuzzy inference system to infer the precise radar PRI information. The core of the fuzzy inference system is the fuzzy rule base, and the fuzzy rule expression of the fuzzy rule base is as follows:
[0039]
[0040] Among them, R l represents the l-th rule in the fuzzy rule base, x PRI represents the radar PRI value measured by the jammer as the input of the fuzzy inference system, and x δ represents the error value between the radar PRI value inferred by the jammer at the previous moment and the current measured value as the input of the fuzzy inference system. represents the precise value of the inferred radar PRI, respectively represent the fuzzy sets to which the measured radar PRI value, the error value between the radar PRI value inferred by the jammer at the previous moment and the current measured value, and the precise value of the inferred radar PRI belong.
[0041] During the description of the fuzzy rules, it is called the antecedent of the rule, that is, the IF part of the rule, and it is called the consequent of the rule, that is, the THEN part of the rule.
[0042] Then, based on the constructed fuzzy rule base, precise inference of the radar PRI is carried out. The overall inference process includes: fuzzification, fuzzy inference, and defuzzification.
[0043] The fuzzification adopts a triangular membership function, and the specific expression is as follows:
[0044]
[0045] Among them, μ tri (x; a, b, n) represents the triangular membership function, x represents the input variable of the fuzzy inference system, a represents the lower bound of the parameter, b represents the upper bound of the parameter, n represents the vertex of the triangle, and a < n < b.
[0046] The specific expression of the fuzzy inference process is as follows:
[0047]
[0048] Among them, represents the activation degree of different input variables of the fuzzy system under rule l, and x PRI,m represents the measured radar PRI, and x δ,m represents the error between the measured value of the radar PRI and the previous inferred value of the radar PRI. and respectively represent the membership functions of the input measured radar PRI and the measurement error, and N r represents the total number of rules in the fuzzy rule base.
[0049] Defuzzification operation is performed and realized by the center-average method. The specific expression is as follows:
[0050]
[0051] Among them, represents the consequent of the rule selected under different rules.
[0052] Based on the interception time t R,m-1 of the (m - 1)-th radar pulse perceived by the jammer and the inferred radar PRI information the state information t R,m required for the jammer to make a decision is determined. The specific expression is as follows:
[0053]
[0054] S32. Establish a generalized Markov decision process for tow jamming and radar countermeasure;
[0055] It is assumed that the jammer is regarded as an agent, and the process of tow jamming and radar countermeasure is modeled as a generalized Markov decision process, which is represented by a quadruple and the specific definition is as follows:
[0056] (1) State set The state is calculated from the time point when the jammer intercepts the (m - 1)-th pulse signal emitted by the radar and the radar PRI refined by the fuzzy inference system, that is, the state is t R,m .
[0057] (2) Action set The action set is the time t J,m when the jammer emits a range forward tow jammer, which depends on the PRI value measured and inferred by the jammer when intercepting the (m - 1)-th radar pulse signal and the time delay of the range forward tow jammer that the m-th pulse should modulate.
[0058] (3) State transition probability T: The state transition probability depends on the dynamic characteristics of the external electromagnetic environment interacting with the jammer.
[0059] (4) Reward value The reward function of the jammer is a function with the number of consecutive pulled radar CPIs as the independent variable, expressed as
[0060] Among them, R m represents the reward value obtained by the jammer for transmitting the m-th interference pulse, and r m (·) represents the designed exponential reward function, with the number of consecutive towed radar pulses as the independent variable.
[0061] Furthermore, the specific steps of step S4 are as follows:
[0062] S41. Obtain the input values of the fuzzy actor-critic algorithm and initialize them;
[0063] Set the intercepted radar pulses, and the measured radar PRI information is x PRI , and the generated inference error is x δ ; The trace decay rate λ θ of the actuator update factor θ ∈ [0, 1], the trace decay rate λ w of the evaluator update factor w ∈ [0, 1], the actuator learning step size α θ > 0, and the evaluator learning step size α w > 0.
[0064] S42. Set a total of I data frames for co-training. Each data frame is regarded as one iteration. The jammer senses the arrival time t R,m-1 of the (m - 1)-th radar pulse, and uses fuzzy inference to obtain the refined radar PRI information to obtain the refined state information
[0065] S43. Select the action t J,m according to the exponential softmax distribution strategy;
[0066] The expression of the exponential softmax distribution strategy π(t J,m |t R,m , θ) is as follows:
[0067]
[0068] Among them, e ≈ 2.71828 is the base of the natural logarithm, and the function h(t R,m , t J,m , θ) represents a parameterized numerical preference, which can be arbitrarily parameterized, θ represents the actuator factor, and t′ J,mRepresents the action selection of the jammer for state t R,m and action t J,m , h(t R , m, t J , m, θ) The larger it is, the higher the probability of selecting action t R,m under state t J,m . For h(t R , m, t J , m, θ), it is represented by a simple linear combination of features, and its expression is as follows:
[0069] h(t R,m , t J,m , θ) = θ T x(t R,m , t J,m ), (17)
[0070] where x(t R,m , t J,m ) represents the feature vector, and (·) T represents the transpose operator.
[0071] S44. The jammer executes action t J,m , that is, the jammer emits a towing jam at time t J,m to obtain the reward value r m+1 , and obtains a new input value;
[0072] S45. Calculate the temporal difference error
[0073] The calculation expression is as follows:
[0074]
[0075] where γ represents the discount factor, and represent the fitted values of the state value function at states t R,m+1 and t R,m , and w represents the evaluator factor. Calculate the evaluator eligibility trace vector The specific expression is as follows:
[0076]
[0077] where, and respectively represent the evaluator eligibility trace vectors when the jammer intercepts the mth radar pulse and the (m + 1)th radar pulse, represents the differential operator, and s m represents the radar pulse state information obtained by the jammer, that is, the radar pulse appearance time t R,m。
[0078] S46. Calculate the actuator eligibility trace vector The specific expression is as follows:
[0079]
[0080] Where, and respectively represent the actuator eligibility trace vectors when the jammer intercepts the m-th radar pulse and the (m + 1)-th radar pulse, represents the gradient update factor of the actuator eligibility trace vector, and the specific expression is as follows:
[0081]
[0082] Update the critic factor w, and the specific expression is as follows:
[0083]
[0084] Where, w m+1 and w m respectively represent the critic factors when the jammer intercepts the (m + 1)-th radar pulse and the m-th radar pulse, and α w represents the learning step size of the critic factor, represents the temporal difference error when the jammer intercepts the (m + 1)-th radar pulse.
[0085] S47. Update the actuator factor θ, and the specific expression is as follows:
[0086]
[0087] Where, θ m+1 and θ m respectively represent the actuator factors when the jammer intercepts the (m + 1)-th radar pulse and the m-th radar pulse, and α θ represents the learning step size of the actuator factor.
[0088] S48. Determine whether the termination state is reached. If the termination state has not been reached, that is, the algorithm has not converged, return to step S43; otherwise, execute step S49;
[0089] S49. Output the emission time of the range pull-off interference pulse;
[0090] The jammer makes a decision on the emission time of the interference pulse in each data frame and outputs the emission time of the range pull-off interference pulse. Finally, the radar range gate is pulled, and the radar working mode is switched from the tracking mode to the search mode.
[0091] Advantages of the present invention: The method of the present invention first establishes a target echo model and an interference signal model according to the confrontation scenario between a radar and a jammer. Regarding the jammer as an agent, a mathematical optimization model for generating a towed interference strategy is established. Then, a fuzzy inference system is used to accurately infer the radar pulse repetition interval sensed by the jammer to obtain the accurate time when the radar pulse arrives at the interference receiver. The process of towed interference and radar confrontation is modeled as a generalized Markov decision process. Finally, the proposed fuzzy actor-critic algorithm is used to solve the problem of generating the towed interference time-domain strategy, and the strategy for the towed interference emission moment is obtained. The method of the present invention can use a fuzzy inference system to accurately determine the towed interference emission moment based on the incomplete information of the radar pulse repetition interval received by the jammer. By modeling the process of towed interference and radar confrontation as a generalized Markov decision process, in the time domain, the moment when the jammer emits towed interference is regarded as an action, and the moment when the radar pulse appears is regarded as a state. The number of consecutive towed radar CPI is used as a reward function to achieve the time-domain deception interference effect and effectively tow the radar gate away from the true target echo signal. Description of the Drawings
[0092] Figure 1 It is a flowchart of a method for generating a towed interference time-domain strategy based on a fuzzy actor-critic algorithm of the present invention.
[0093] Figure 2 It is a confrontation scenario diagram between a jammer and a radar in an embodiment of the present invention.
[0094] Figure 3 It is a time-domain JSR result diagram under simulation of multiple methods in an embodiment of the present invention.
[0095] Figure 4 It is a result diagram of the number of time-domain radar CPI interfered under simulation of multiple methods in an embodiment of the present invention.
[0096] Figure 5 It is a result diagram of the probability of time-domain radar CPI being interfered under simulation of multiple methods in an embodiment of the present invention.
[0097] Figure 6 It is a result diagram of the number of consecutive towed data frames under simulation of multiple methods in an embodiment of the present invention.
[0098] Figure 7 It is a result diagram of the distance between the continuously towed range gate under simulation of multiple methods in an embodiment of the present invention. Detailed Embodiment
[0099] The method of the present invention will be further described below in conjunction with the drawings and embodiments.
[0100] As Figure 1As shown in the figure, the flowchart of a method for generating a time-domain strategy for towing interference based on a fuzzy actor-critic algorithm according to the present invention is as follows:
[0101] S1. Establish a target echo model and an interference signal model according to the confrontation scenario between the radar and the jammer;
[0102] S2. Establish a mathematical optimization model for generating a towing interference strategy based on the mechanism of towing interference and radar confrontation;
[0103] S3. Regard the jammer as an agent, use a fuzzy inference system to precisely process the incomplete information of the radar PRI perceived by the jammer, obtain the time when the radar pulse arrives at the receiver as the state, and establish a generalized Markov decision process for towing interference and radar confrontation;
[0104] S4. Use the fuzzy actor-critic algorithm to solve the established generalized Markov decision problem, obtain the strategy selection for the towing interference emission time, and ensure that the jammer actively towes the radar range gate under the condition of incomplete radar information acquisition.
[0105] In this embodiment, the specific steps of step S1 are as follows:
[0106] During the confrontation process, the radar emits a pulse train to irradiate the target, and the total number of tracking frames is M, which is continuously repeated at the pulse repetition interval PRI (Pulse repetition interval). The confrontation process between the jammer and the radar is as Figure 2 shown. Then the m-th pulse signal s c (t) emitted by the radar at the center frequency f m has the following expression:
[0107] s m (t) = A R s(t - mT P,R,m ) exp(j2πf c t) (1)
[0108] where m = 1, 2,..., M; t represents time, A R represents the amplitude of the radar emission signal, T P,R,m represents the PRI of the m-th pulse, and s(t) represents a linear frequency-modulated signal LFM (Linear frequency-modulated). The specific waveform expression is as follows:
[0109]
[0110] where T d,R represents the pulse width, μ = B / T d,R represents the frequency modulation slope, B represents the signal bandwidth, and rect(·) represents the rectangular function, and the expression is as follows:
[0111]
[0112] Due to the influence of the distance between the jammer and the radar, the LFM signal will generate a time delay t during propagation. R = R(t) / c, where R(t) is affected by the movement of the jammer and the radar and is time-related. Due to the non-cooperative relationship between the jammer and the radar, the jammer does not know the moment when the radar emits a pulse, but can sense the moment when the radar pulse signal is intercepted by the jammer. When the jammer intercepts the m-th pulse signal x m (t) of the radar, the expression is as follows:
[0113] x m (t) = α m s m (t - t R ), (4)
[0114] where α m represents the propagation path loss, t R = R(t) / c represents the time delay generated by the LFM signal during propagation, R(t) represents the distance function between the target and the radar, and c = 3×10 8 (m / s) represents the speed of light.
[0115] Then the jammer measures the parameters of the radar pulse signal. The parameter measurement includes: the center frequency of the radar signal, the signal bandwidth, direction finding and angle measurement of the radar position.
[0116] By sensing the radar pulse emission law, before the radar pulse reaches the jammer receiver, the jammer emits a range forward towing interference signal in advance and controls the forwarding time delay to gradually decrease, simulating the target approaching the radar to exert pressure on the radar side. Then the jammer emits a range towing interference signal J m (t J,m ) with the following expression:
[0117] J m (t J,m ) = A J α m s m (t J,m - t R + Δt J (t)) (5)
[0118] where A J represents the amplitude of the range towing interference signal emitted by the jammer, t J,m represents the emission moment of the m-th pulse of the range forward towing interference. And in order to effectively tow the radar range gate, there is A J > A R, Δt J (t) represents the forward delay function of the distance from the traction interference.
[0119] Set v J represents the speed during uniform traction, then the forward delay function Δt of the distance from the traction interference J (t) is expressed as follows:
[0120]
[0121] Among them, T J represents the total time of one traction interference, which is divided into three parts according to the three stages of the traction interference. t1 represents the cut-off moment of the interference capture period, t2 represents the cut-off moment of the interference traction period. From Figure 2 it can be known that since the jammer modulates a time delay for the intercepted radar signal, the distance of the false target detected by the radar is ΔR J,m .
[0122] In this embodiment, the step S2 is specifically as follows:
[0123] To ensure the effectiveness of the distance traction interference, the interference pulse needs to continuously pull the range gate. When the interference pulls the range gate, due to the width setting of the range gate, there will be a period of time when both the radar pulse and the interference pulse are within the range gate. At this time, once an interference pulse fails to pull the range gate normally, that is, the traction is interrupted. Due to the certain memory ability of the range gate, the range gate will re-capture the real target echo pulse, resulting in the invalidation of the previous interference pulse traction. Therefore, the distance traction interference needs to maximize the continuous traction of the radar CPI number until the range gate is completely pulled away from the real target echo signal, and the radar switches from the tracking mode to the search mode, and the forward distance traction interference is successful.
[0124] The mathematical optimization model expression for generating the traction interference strategy is as follows:
[0125]
[0126] Among them, t J,m represents the transmission moment of the m-th distance traction interference pulse, represents the deduced m-th radar PRI, t min ≤t J,m ≤t max represents all interference pulses received within the defined time window of the radar, t min and t max represent the minimum and maximum values of the time window range respectively; represents the specific range of the jammer's deduced radar PRI, T min represents the minimum value of the jammer's deduced radar PRI, T maxRepresents the maximum value of the radar PRI inferred by the jammer; Represents the number of CPI of the continuous towing target echo of the range forward towing jammer, and the specific expression is as follows:
[0127]
[0128] Among them, Represents the time-domain overlap probability between the range forward towing jammer pulse and the radar pulse, τ represents the probability threshold, when it indicates that the range forward towing jammer can accurately infer the radar signal PRI information to perform forward towing on each radar pulse, and its specific expression is as follows:
[0129]
[0130] Among them, Represents the time-domain range of the mth pulse of the radar, t R,m-1 Represents the time when the jammer intercepts the (m - 1)th pulse of the radar, Represents the time-domain range of the mth pulse of the range forward towing jammer, T d,J Represents the width of the jammer pulse, then t J,m The specific expression is as follows:
[0131]
[0132] In this embodiment, the step S3 is specifically as follows:
[0133] S31. Use a fuzzy inference system to precisely process the incomplete information of the radar PRI perceived by the jammer;
[0134] Adopt a fuzzy inference system FIS with multiple inputs - single output MISO. Utilize the function approximation ability of the fuzzy inference system to infer the accurate radar PRI information. The core of the fuzzy inference system is the fuzzy rule base, and the fuzzy rule expression of the fuzzy rule base is as follows:
[0135]
[0136] Among them, R l Represents the lth rule in the fuzzy rule base, x PRI Represents the radar PRI value measured by the jammer as the input of the fuzzy inference system, x δ Represents the error value between the radar PRI value inferred by the jammer at the previous moment and the current measured value as the input of the fuzzy inference system, Represents the accurate value of the inferred radar PRI, They respectively represent the measured radar PRI value, the error value between the radar PRI value inferred by the jammer at the previous moment and the current measured value, and the fuzzy set to which the accurately inferred radar PRI value belongs.
[0137] During the description of fuzzy rules, is called the antecedent of the rule, that is, the IF part of the rule, and is called the consequent of the rule, that is, the THEN part of the rule.
[0138] Then, based on the constructed fuzzy rule base, precise inference is carried out on the radar PRI. The overall inference process includes: fuzzification, fuzzy inference, and defuzzification.
[0139] For the sake of simplicity in calculation and reduction of the amount of calculation, triangular membership functions are adopted for fuzzification. The specific expression is as follows:
[0140]
[0141] Among them, μ tri (x; a, b, n) represents the triangular membership function, x represents the input variable of the fuzzy inference system, a represents the lower bound of the parameter, b represents the upper bound of the parameter, n represents the vertex of the triangle, and a < n < b.
[0142] The fuzzy inference process is based on the fuzzy rule base. After obtaining the membership functions of the two input variables of the measured radar PRI and the measurement error, logical operations are performed using fuzzy operators to obtain the activation degree of each rule. The specific expression of the fuzzy inference process is as follows:
[0143]
[0144] Among them, represents the activation degree of different input variables of the fuzzy system under rule l, x PRI,m represents the measured radar PRI, x δ,m represents the error between the measured radar PRI value and the previously inferred radar PRI value, and respectively represent the membership functions of the input radar PRI measurement value and the measurement error, and N r represents the total number of rules in the fuzzy rule base.
[0145] After the fuzzy inference process obtains the activation degree of each rule, the output value is still not a specific value. To determine the final output value of the fuzzy inference system, defuzzification operation is required. In this embodiment, the defuzzification operation is implemented by the center average method. The specific expression is as follows:
[0146]
[0147] Among them, Represents the consequent of the rule selected under different rules.
[0148] Based on the interception time t of the (m - 1)-th radar pulse sensed by the jammer R,m-1 and the deduced radar PRI information Determine the state information t required for the jammer to make a decision. R,m , and the specific expression is as follows:
[0149]
[0150] S32. Establish a generalized Markov decision process for towing jamming and radar countermeasure;
[0151] Set the jammer as an agent, and model the process of towing jamming and radar countermeasure as a generalized Markov decision process, which is represented by a quadruple , and the specific definition is as follows:
[0152] (1) State set The state is calculated from the time point when the jammer intercepts the (m - 1)-th pulse signal transmitted by the radar and the radar PRI refined by the fuzzy inference system, that is, the state is t R,m .
[0153] (2) Action set The action set is the time t when the jammer emits a range forward towing jammer J,m , which depends on the PRI value measured and deduced by the jammer when intercepting the (m - 1)-th radar pulse signal and the time delay of the range forward towing jammer that the m-th pulse should modulate.
[0154] (3) State transition probability T: The state transition probability depends on the dynamic characteristics of the external electromagnetic environment interacting with the jammer.
[0155] (4) Reward value The reward function of the jammer is a function with the continuous towing radar CPI number as the independent variable, which is expressed as
[0156] Among them, R m Represents the reward value obtained by the jammer when emitting the m-th jammer pulse, and r m (·) represents the designed exponential reward function, with the continuous towing radar pulse number as the independent variable.
[0157] In this embodiment, the specific steps of step S4 are as follows:
[0158] S41. Obtain the input values of the fuzzy actor-critic algorithm and initialize them;
[0159] Set the intercepted radar pulses, and the measured radar PRI information is x PRI , and the generated inference error is x δ ; the trace decay rate λ of the actuator update factor θ θ ∈[0, 1], the trace decay rate λ of the evaluator update factor w w ∈[0, 1], the actuator learning step size α θ > 0, the evaluator learning step size α w > 0.
[0160] S42. Set a total of I data frames for co-training. Each data frame is regarded as one iteration. The jammer senses the occurrence time t of the (m - 1)-th radar pulse R,m-1 , and uses fuzzy inference to obtain the refined radar PRI information Obtain the refined state information
[0161] S43. Select the action t according to the exponential softmax distribution strategy J,m ;
[0162] The exponential softmax distribution strategy π(t J,m |t R,m , θ) is expressed as follows:
[0163]
[0164] where e≈2.71828 is the base of the natural logarithm, and the function h(t R,m , t J,m , θ) represents a parameterized numerical preference, which can be parameterized arbitrarily. θ represents the actuator factor, and t′ J,m represents the action selection of the jammer. For the state t R,m and the action t J,m , the larger h(t R , m, t J , m, θ) is, the higher the probability of selecting the action t R,m in the state t J,m . For h(t R , m, t J , m, θ), it is represented by a simple linear combination of features, and its expression is as follows:
[0165] h(t R,m , t J,m , θ) = θ T x(t R,m , t J,m ), (17)
[0166] where x(t R,m , t J,m) represents the feature vector, (·) T represents the transpose operator.
[0167] S44. The jammer executes action t J,m , that is, the jammer emits a towing jam at time t J,m and obtains a reward value r m+1 , and acquires a new input value;
[0168] S45. Calculate the temporal difference error
[0169] The calculation expression is as follows:
[0170]
[0171] where γ represents the discount factor, and represent the fitted values of the state value function at states t R,m+1 and t R,m respectively, and w represents the evaluator factor. Calculate the evaluator eligibility trace vector The specific expression is as follows:
[0172]
[0173] where, and represent the evaluator eligibility trace vectors when the jammer intercepts the m - th radar pulse and the (m + 1)-th radar pulse respectively, represents the differential operator, and s m represents the radar pulse state information obtained by the jammer, that is, the occurrence time t of the radar pulse R,m .
[0174] S46. Calculate the actuator eligibility trace vector The specific expression is as follows:
[0175]
[0176] where, and represent the actuator eligibility trace vectors when the jammer intercepts the m - th radar pulse and the (m + 1)-th radar pulse respectively, represents the gradient update factor of the actuator eligibility trace vector, and the specific expression is as follows:
[0177]
[0178] Update the evaluator factor w, and the specific expression is as follows:
[0179]
[0180] Among them, w m+1 and w m represent the evaluator factors when the (m + 1)-th radar pulse and the m-th radar pulse are intercepted by the jammer respectively, and α w represents the learning step size of the evaluator factor, represents the temporal difference error when the (m + 1)-th radar pulse is intercepted by the jammer.
[0181] S47. Update the actuator factor θ, and the specific expression is as follows:
[0182]
[0183] Among them, θ m+1 and θ m represent the actuator factors when the (m + 1)-th radar pulse and the m-th radar pulse are intercepted by the jammer respectively, and α θ represents the learning step size of the actuator factor.
[0184] S48. Judge whether the termination state is reached. If the termination state has not been reached, that is, the algorithm has not converged, return to step S43; otherwise, execute step S49.
[0185] S49. Output the emission time of the range towing interference pulse;
[0186] The jammer makes a decision on the emission time of the interference pulse in each data frame and outputs the emission time of the range towing interference pulse. Finally, the range gate of the radar is towed, and the radar working mode is converted from the tracking mode to the search mode.
[0187] This embodiment further conducts simulation verification and analysis on the method of the present invention, as follows:
[0188] When the radar is in the tracking mode and the jammer emits self-defense range towing interference, the results within 700 data frames are presented as a whole. There are 30 coherent processing intervals in each data frame. The radar transmission signal power is 5 kW, the radar transmitting antenna gain and the receiving antenna gain are both 30 dB, the radar range gate length is set to 350 m, the radial distance between the radar and the target is 10 km, and the target radar cross section is 1 m 2 , the jammer transmission power is 10 kW, and the towing speed is 250 m / s.
[0189] The basic parameter settings of the fuzzy actor-critic algorithm are as follows: the learning step size α w of the evaluator is initially set to 0.4, the learning step size α θ of the actuator is initially set to 0.2, and the learning step size α wand the actuator learning step size α θ Both will gradually decrease as the algorithm executes; the discount rate γ is set to 0.9. In all simulation results, the curve graphs and bar graphs are the results after Monte Carlo. The hardware and software operating conditions for the digital simulation experiment are as follows: the operating system is Windows 11, the Intel Core i5-12400 CPU is 4.40GHz, the RAM is 32.0GB, and the operating software is MATLAB R2023a.
[0190] To fully verify the effectiveness of the fuzzy actor-critic algorithm (TSFAC), three typical strategy generation algorithms are selected for comparison:
[0191] (1) Use the passive algorithm for traction interference time resource scheduling (PASS): The jammer intercepts the radar pulse, completes parameter measurement, and then passively forwards the signal after waveform modulation.
[0192] (2) Use the active fixed algorithm for jammer interference time resource scheduling (AFTSS): The jammer uses a pre-constructed strategy library, selects a template matching strategy based on the radar pulse intercepted at the previous moment, and selects a fixed emission strategy after waveform modulation.
[0193] (3) Use the active Q-learning algorithm for jammer interference time resource scheduling (ATSSQL): The jammer uses the Q-learning algorithm to generate a strategy for the traction interference emission moment.
[0194] Figure 3 This is the time-domain JSR result graph under multiple simulation methods in the embodiment of the present invention. From Figure 3 It can be seen that the TSFAC algorithm proposed by the method of the present invention has the optimal steady-state performance and a relatively fast convergence speed. The TSFAC algorithm uses a fuzzy inference system to precisely process the measured radar PRI information, determines the radar pulse emission moment as the state, makes decisions by fully utilizing prior knowledge and expert knowledge, and demonstrates good interference effectiveness. Compared with the ATSSQL algorithm, the JSR is increased by about 1.5dB, compared with the PASS algorithm, the JSR is increased by about 2.5dB, and compared with the AFTSS algorithm, the JSR is increased by more than 3dB. This shows that the strategy generated by the proposed TSFAC algorithm can significantly improve the interference effectiveness of the jammer in a complex electromagnetic environment and has strong interference capabilities.
[0195] Figure 4 This is the time-domain radar CPI number of interfered results graph under multiple simulation methods in the embodiment of the present invention. From Figure 4It can be seen that under different interference emission time strategies, the number of interfered CPIs of the radar shows obvious differences. Within a single data frame, the number of interfered CPIs using the AFTSS algorithm is only below 14. After convergence, the number of interfered CPIs of other algorithms can all reach above 20. The TSFAC algorithm proposed by the method of the present invention still has the optimal performance.
[0196] Figure 5 This is the result graph of the probability of the time-domain radar CPI being interfered under multiple simulation methods in the embodiment of the present invention. The TSFAC algorithm proposed by the method of the present invention can use the fuzzy inference system to accurately perceive and reason about the state of the radar transmitted signal, and has the ability to process high-dimensional state spaces, and can generate continuous-time actions during the time-domain resource scheduling process. Although the ATSSQL algorithm also shows certain superiority, its ability to process high-dimensional action spaces under uncertain information states is still insufficient. It can only be processed after dimensionality reduction through state space discretization operations, which will generate certain discretization errors.
[0197] Figure 6 This is the result graph of the number of consecutive towed data frames under multiple simulation methods in the embodiment of the present invention. Figure 7 This is the result graph of the distance of the continuous towed range gate under multiple simulation methods in the embodiment of the present invention. It can be seen that the number of consecutive interfering data frames will affect the towing distance of the towed interference to the range gate. Among them, the TSFAC algorithm proposed by the method of the present invention can have the largest number of consecutive interfering data frames and the longest towing distance, and can successfully tow the range gate away from the real target echo signal, forcing the radar to switch from the tracking mode to the search mode. Although the other interference strategies can produce a certain towing effect on the range gate, they cannot successfully tow the range gate away from the real target echo signal. This also means that the other interference strategies are all failed during the process of releasing the towed interference.
[0198] In summary, the method of the present invention can use the fuzzy inference system for precise processing according to the incomplete information of the radar PRI intercepted at the current moment, obtain the radar pulse arrival time as the state, and by modeling the radar countermeasure process of the jammer as a generalized Markov decision process, regarding the interference emission time of the jammer in the time domain as an action and the radar pulse appearance time as the state, and using the number of consecutive towed CPIs of the towed interference as the reward function, to achieve the time-domain deception interference effect and effectively tow the radar gate away from the real target echo signal.
[0199] Those of ordinary skill in the art will realize that the embodiments described herein are provided to assist the reader in understanding the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on these technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.
Claims
1. A method for generating a time-domain strategy for towing interference based on a fuzzy actor-critic algorithm, the specific steps are as follows: S1. According to the confrontation scenario between the radar and the jammer, establish a target echo model and a jamming signal model; S2. Based on the mechanism of towing interference and radar confrontation, establish a mathematical optimization model for generating towing interference strategies; S3. Regard the jammer as an agent, use a fuzzy inference system to precisely process the incomplete information of the radar PRI perceived by the jammer, obtain the time when the radar pulse arrives at the receiver as the state, and establish a generalized Markov decision process for towing interference and radar confrontation; S4. Use the fuzzy actor-critic algorithm to solve the established generalized Markov decision problem, obtain the strategy selection of the towing interference emission time, and ensure that the jammer actively towes the radar range gate under the condition of incomplete radar information acquisition.
2. The method for generating a traction interference time-domain strategy based on the fuzzy actor-critic algorithm according to claim 1, wherein The specific steps of step S1 are as follows: During the confrontation process, the radar transmits a pulse train to illuminate the target. The total number of tracking frames is M, which is repeated at the pulse repetition interval PRI. The radar uses the center frequency f c The mth pulse signal transmitted s m (t) is expressed as follows: s m (t) = A R s(t - mT P,R,m ) exp(j2πf c t) (1) where m = 1, 2, …, M; t represents time, A R represents the amplitude of the radar transmitted signal, T P,R,m represents the PRI of the m-th pulse, and s(t) represents the linear frequency modulation signal LFM. The specific waveform expression is as follows: Among them, T d,R represents the pulse width, μ = B / T d,R represents the frequency modulation slope, B represents the signal bandwidth, rect(·) represents the rectangular function, and the expression is as follows: The jammer can sense the moment when the radar pulse signal is intercepted by the jammer; when the jammer intercepts the m-th radar pulse signal x m (t) is expressed as follows: x m (t) = α m s m (t - t R ), (4) where α m represents the propagation path loss, and t R = R(t) / c represents the time delay generated by the LFM signal during propagation, R(t) represents the distance function between the target and the radar, and c represents the speed of light; Then, the jammer measures the parameters of the radar pulse signal, and the parameter measurement includes: the center frequency of the radar signal, the signal bandwidth, direction finding and angle measurement of the radar position; By sensing the radar pulse emission pattern with a jammer, before the radar pulse reaches the jammer receiver, a range forward-dragging interference signal is transmitted in advance, and the forwarding delay is controlled to gradually decrease, simulating a target approaching the radar to exert pressure on the radar side. Then the jammer transmits a range-dragging interference signal J m (t J,m ) The expression is as follows: J m (t J,m ) = A J α m s m (t J,m -t R +Δt J (t)) (5) Among them, A J represents the amplitude of the distance towing interference signal emitted by the jammer, and t J,m represents the emission time of the m-th pulse of the forward towing interference, and A J > A R , and Δt J (t) represents the forward delay function of the distance towing interference; Set v J Represents the speed during uniform traction, and the forwarding delay function Δt J (t) has the following expression: Among them, T J represents the total time of a single traction interference, which is divided into three parts according to the three stages of the traction interference. t1 represents the cut-off moment of the interference capture period, and t2 represents the cut-off moment of the interference traction period.
3. A method for generating a traction interference time-domain strategy based on a fuzzy actor-critic algorithm according to claim 1, characterized in that The specific steps of step S2 are as follows: The expression of the mathematical optimization model for generating towing interference strategies is as follows: Among them, t J,m represents the transmission time of the m-th range tow interference pulse, represents the deduced m-th radar PRI, t min ≤t J,m ≤t max represents all interference pulses received within the defined time window of the radar, t min and t max represent the minimum and maximum values of the time window range respectively; represents the specific range of the jammer's deduced radar PRI, T min represents the minimum value of the jammer's deduced radar PRI, T max represents the maximum value of the jammer's deduced radar PRI; represents the number of coherent processing intervals CPI of the range forward tow interference continuous tow target echo, and the specific expression is as follows: Among them, represents the time-domain overlapping probability between the forward towing interference pulse and the radar pulse, τ represents the probability threshold. When it indicates that the forward towing interference in distance can accurately infer the PRI information of the radar signal, and its specific expression is as follows: Among them, represents the time domain range of the m-th pulse of the radar, and t R,m-1 represents the moment when the jammer intercepts the (m - 1)-th pulse of the radar. represents the time domain range of the m-th pulse of the forward towing jammer, and T d,J represents the width of the interference pulse, then t J,m The specific expression is as follows:
4. A method for generating a traction interference time-domain strategy based on a fuzzy actor-critic algorithm according to claim 1, characterized in that The specific steps of step S3 are as follows: S31. Use a fuzzy inference system to precisely process the incomplete information of the radar PRI perceived by the jammer; Adopt a multi-input single-output (MISO) fuzzy inference system FIS, use the function approximation ability of the fuzzy inference system to infer the accurate radar PRI information. The core of the fuzzy inference system is the fuzzy rule base, and the fuzzy rule expression of the fuzzy rule base is as follows: where, R l represents the l-th rule in the fuzzy rule base, and x PRI represents the radar PRI value measured by the jammer as the input of the fuzzy inference system, and x δ represents the error value between the radar PRI value inferred by the jammer at the previous moment and the current measured value as the input of the fuzzy inference system, represents the inferred exact value of the radar PRI, respectively represent the fuzzy sets to which the measured radar PRI value, the error value between the radar PRI value inferred by the jammer at the previous moment and the current measured value, and the inferred exact value of the radar PRI belong; In the process of fuzzy rule description, it is called the antecedent of the rule, that is, the IF part of the rule, and it is called the consequent of the rule, that is, the THEN part of the rule. Then, based on the constructed fuzzy rule base, perform accurate inference on the radar PRI. The overall inference process includes: fuzzification, fuzzy inference, and defuzzification; The fuzzification adopts a triangular membership function, and the specific expression is as follows: Among them, μ tri (x; a, b, n) represents a triangular membership function, where x represents the input variable of the fuzzy inference system, a represents the lower bound of the parameter, b represents the upper bound of the parameter, n represents the vertex of the triangle, and a < n < b; The specific expression of the fuzzy inference process is as follows: Among them, represents the activation degree of different input variables of the fuzzy system under rule l, x PRI,m represents the measured radar PRI, x δ,m represents the error between the measured value of the radar PRI and the previous inferred value of the radar PRI, and respectively represent the membership functions of the input radar PRI measured value and the measurement error, N r represents the total number of rules in the fuzzy rule base; Perform defuzzification operations, which are realized by the center average method, and the specific expression is as follows: Among them, represents the consequent of the rule selected under different rules; Based on the interception time t of the (m - 1)-th radar pulse sensed by the jammer R,m-1 and the deduced radar PRI information determine the state information t required for the jammer to make a decision R,m , and the specific expression is as follows: S32. Establish a generalized Markov decision process for towing interference and radar confrontation; It is assumed that the jammer is regarded as an agent, and the process of towing jamming and radar countermeasure is modeled as a generalized Markov decision process, which is represented by a quadruple as follows: (1) State set The state is calculated from the time point when the jammer intercepts the (m - 1)-th pulse signal transmitted by the radar and the radar PRI refined by the fuzzy inference system, that is, the state is t R,m ; (2) Action set The action set is the moment t when the jammer emits range forward towing jamming J,m , which depends on the PRI value measured and inferred from the (m - 1)-th radar pulse signal intercepted by the jammer and the range forward towing jamming time delay that the m-th pulse should be modulated (3) State transition probability T: The state transition probability depends on the dynamic characteristics of the external electromagnetic environment interacting with the jammer; (4) Reward value The reward function of the jammer is a function with the number of consecutive coherent processing intervals (CPI) of the towed radar as the independent variable, expressed as Among them, R m represents the reward value obtained when the jammer emits the m-th interference pulse, and r m (·) represents the designed exponential reward function, with the number of consecutive towing radar pulses as the independent variable.
5. A method for generating a traction interference time-domain strategy based on a fuzzy actor-critic algorithm according to claim 1, characterized in that, The specific steps of step S4 are as follows: S41. Obtain the input values of the fuzzy actor-critic algorithm and initialize them; Set the intercepted radar pulse, and the measured radar PRI information is x PRI , and the generated inference error is x δ ; the trace decay rate λ of the actuator update factor θ θ ∈ [0, 1], the trace decay rate λ of the evaluator update factor w w ∈ [0, 1], the actuator learning step size α θ > 0, the evaluator learning step size α w > 0; S42. Set a total of I data frames for co-training. Each data frame is regarded as one iteration. The jammer senses the occurrence time t of the (m - 1)-th radar pulse R,m-1 , and uses fuzzy inference to obtain refined radar PRI information Obtain refined state information S43. Select an action t according to the exponential flexible maximization distribution strategy J,m ; Exponential Flexible Maximal Distribution Strategy π(t J,m |t R,m , θ) is expressed as follows: where \(e\approx2.71828\) is the base of the natural logarithm, and the function \(h(t\) R,m ,t J,m ,\theta)\) represents a parameterized numerical preference, which can be arbitrarily parameterized, \(\theta\) represents the actuator factor, and \(t'\) J,m represents the action selection of the jammer. For the state \(t\) R,m and the action \(t\) J,m , the larger \(h(t\) R,m ,t J,m ,\theta)\) is, the higher the probability of selecting the action \(t\) R,m in the state \(t\) J,m . For \(h(t\) R,m ,t J,m ,\theta)\), it is represented by a simple linear combination of features, and its expression is as follows: h(t R,m ,t J,m ,θ)=θ T x(t R,m ,t J,m ), (17) Among them, x(t R,m , t J,m ) represents the feature vector, and (·) T represents the transpose operator; S44. The jammer performs action t J,m , that is, the jammer emits a towing jam at time t J,m to obtain a reward value r m+1 , and obtains a new input value; S45. Calculate the temporal difference error The calculation expression is as follows: where γ represents the discount factor, and represents the fitted value of the state-value function at state t R,m+1 and t R,m The fitted value of the state-value function, w represents the critic factor; calculate the critic eligibility trace vector The specific expression is as follows: Among them, and respectively represent the evaluator eligibility trace vectors when the m-th radar pulse and the (m + 1)-th radar pulse are intercepted by the jammer. ▽ represents the differential operator, and s m represents the radar pulse state information obtained by the jammer, that is, the radar pulse appearance time t R,m ; S46. Calculate the actuator eligibility trace vector The specific expression is as follows: Among them, and respectively represent the actuator eligibility trace vectors when the jammer intercepts the m-th radar pulse and the (m + 1)-th radar pulse. ▽lnπ(a m |s m , θ) represents the gradient update factor of the actuator eligibility trace vector, and the specific expression is as follows: Update the evaluator factor w, and the specific expression is as follows: where, w m+1 and w m represent the evaluator factors when the (m + 1)-th radar pulse and the m-th radar pulse are intercepted by the jammer, respectively, and α w represents the learning step size of the evaluator factor, represents the temporal difference error when the (m + 1)-th radar pulse is intercepted by the jammer; S47. Update the actuator factor θ, and the specific expression is as follows: Among them, θ m+1 and θ m represent the actuator factors when the (m + 1)-th radar pulse and the m-th radar pulse are intercepted by the jammer respectively, and α θ represents the learning step size of the actuator factor; S48. Determine whether the termination state is reached. If the termination state has not been reached, that is, the algorithm has not converged, return to step S43; otherwise, execute step S49; S49. Output the time of the towing interference pulse emission; The jammer makes a decision on the interference pulse emission time in each data frame and outputs the time of the towing interference pulse emission; finally, realize the towing of the radar range gate, and convert the radar working mode from the tracking mode to the search mode.