Networking radar distributed cooperative track deception jamming resisting method based on reinforcement learning

By inserting a review time slot into the network radar and optimizing multi-domain parameter scheduling using reinforcement learning algorithms, the problem of network radar being difficult to identify distributed collaborative track spoofing interference is solved, efficient target recognition and interference identification are achieved, and the system detection efficiency and resource utilization are improved.

CN120334864APending Publication Date: 2025-07-18UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510538927.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing networking radars are difficult to identify distributed coordinated track fraud interference, especially multi-machine coordinated track fraud interference, resulting in a decrease in situational awareness reliability.

Method used

Using a method based on reinforcement learning, a review time slot is inserted into the network radar, and a review time frame and a review radar are dynamically selected through a composite strategy. Combined with multi-domain resource scheduling, a closed-loop detection mechanism is built, and a reinforcement learning algorithm is used to optimize the multi-domain parameter scheduling strategy to perform target identification and interference identification.

Benefits of technology

Effectively identifying distributed multi-machine coordinated track spoofing interference has improved system detection efficiency and resource utilization, and improved the anti-interference ability of networking radar in complex multi-target scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120334864A_ABST
    Figure CN120334864A_ABST
Patent Text Reader

Abstract

The invention discloses a networking radar distributed cooperative track deception jamming resisting method based on reinforcement learning, and the method comprises the steps: firstly inserting a re-check time slot in a conventional tracking process, transmitting a pulse string which has multi-domain agility after reinforcement learning scheduling in the re-check time slot, and achieving the distributed cooperative track deception jamming in the presence of track deception. A real target is tracked, track cheating of a false target can be identified, then a re-checking time frame and a re-checking radar are dynamically selected through a composite strategy, multi-domain resource scheduling within residence time is combined, the target is checked, finally, whether re-checking is needed or not is judged based on binary hypothesis testing, and construction of a closed-loop detection mechanism is completed. According to the method, the problem that an existing networking radar is difficult to identify such deception jamming is solved, the provided re-checking mechanism can effectively identify distributed multi-aircraft cooperative track deception jamming in a complex multi-target scene, and the system detection efficiency and the resource utilization rate are further improved based on intelligent decision-making of reinforcement learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of networked radars, and particularly relates to a method for anti-distributed cooperative track deception interference of networked radars based on reinforcement learning. Background Technique

[0002] A networked radar is a radar system that integrates multiple radars to work together and has received extensive attention in recent years. Compared with the existing single-radar system, it provides enhanced detection, tracking, and anti-interference capabilities. Supported by data fusion and resource allocation technologies, networked radars have broad development space in the civilian field. However, the emergence of new distributed cooperative track deception interference technologies poses a severe challenge to networked radars: multiple jammers equipped with digital radio frequency memory (DRFM) can, through precise spatio-temporal synchronization, generate false echoes with spatial correlation and temporal consistency at the receiving ends of each node of the networked radar, and then form continuous false tracks that conform to the target motion law at the information fusion center. This type of interference, by simulating the spatial distribution characteristics and motion trajectories of real target groups, makes it difficult for existing anti-interference means based on single-node signal detection or multi-dimensional feature analysis to distinguish, posing a challenge to the reliability of the situation awareness of the networked radar system. The literature "Zhang Xiangyu, Yu Hongbo, Wang Guohong. Radar Networking Technology. Beijing: National Defense Industry Press, 2022" once mentioned that multi-aircraft cooperative interference can pre-design the flight paths of each jammer in advance based on parameters such as the positions of the networked radars that are previously mastered, and intercept radar signals and generate interference signals through cooperative control and the electronic jamming equipment carried by the jammers, so as to achieve false targets with spatial coincidence. At present, the research on anti-deception interference of networked radars mainly focuses on the research of identification technologies for dense false targets around real targets. However, these false targets lack spatial correlation and cannot form the distance-angle joint deception effect of multi-aircraft cooperative track deception interference.

[0003] Distributed multi-aircraft cooperative track deception interference using small-sized unmanned aerial vehicles working in cooperation, because each jammer is located in the airspace close to the radar network, has the characteristics of being closer to the jammed radar, simple equipment, strong survivability, low cost, flexible mobility, small radar cross-section area, being easy to hide and disguise, etc. According to the parameters and position information of the networked radar previously obtained through reconnaissance, through reasonable design, each jammer can be controlled and cooperate to interfere with the networked radar, and the interference signals sent by each jammer can enter from the main lobe of the radar, forming continuous and stable false deception tracks in the networked radar, which has become a new and difficult-to-deal-with interference method in networked radars.

[0004] In the prior art, the document "Suppression of deception-false-target jamming for active / passive netted radar based on position error, IEEE Sens. J., vol. 22, no. 8, pp. 7902-7912, Apr. 2022" uses site error recognition for multi-aircraft cooperative deception jamming; the document "Recognition of track deception jamming based on joint mean-variance test, Acta Aeronautica et Astronautica Sinica, 2016, 42(9): 1680-1685" recognizes multi-aircraft cooperative deception jamming based on radar site detection error and random measurement error. With the development of reconnaissance and jamming technologies, the errors in the reconnaissance or jamming process have been gradually reduced, and the above methods for recognizing multi-aircraft cooperative deception jamming based on errors or deviations have gradually become ineffective. In the research on the anti-deception jamming mechanism of existing networked radars, the document "Cognitive FDA-MIMO radar network for target discrimination and tracking with main-lobe deceptive trajectory interference, IEEE Trans. Aerosp. Electron. Syst., vol. 59, no. 4, pp. 4207-4222, Aug. 2023" uses homologous test based on distance and azimuth differences and the document "Joint resource optimization for a distributed MIMO radar when tracking multiple targets in the presence of deception jamming, Signal Process, vol. 200, Nov. 2022, Art. no. 108641" uses the method of recognizing networked deception jamming based on the spatial distribution of site addresses can only recognize false targets without spatial correlation and cannot be applied to the accurate recognition of distributed multi-aircraft cooperative track deception jamming. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a method for a networked radar to resist distributed cooperative track deception jamming based on reinforcement learning, providing a method for a networked radar to resist advanced distributed cooperative jamming.

[0006] The technical solution adopted by the present invention is as follows: A method for a networked radar to resist distributed cooperative track deception jamming based on reinforcement learning, and the specific steps are as follows:

[0007] S1. In the scenario of distributed multi - machine collaborative deception jamming, obtain the parameter settings of the networked radar system at a certain moment, and initialize the algorithm module according to the training parameters and the parameter settings of the networked radar system;

[0008] S2. The target forms a track at the fusion center, and a review time slot is inserted during the tracking process. The fusion center allocates resources to track the target;

[0009] For the target that forms a track, a review time slot is inserted during the tracking process. That is, for any target q that forms a track at the networked radar fusion center, after the fusion center allocates resources to track this target, it dynamically selects a review time frame as the inserted review time slot, that is, according to the pre - set review interval n c , the review time frame of the q - th target is randomly selected within [n q +b·n c +1,n q +(b + 1)·n c . Before the review time frame , normally track the target and use the track of the fusion center to train the algorithm module corresponding to the q - th target. The algorithm module is denoted as U m,q , m = 1,...M.

[0010] Among them, M represents the number of radar points in the radar network. b = {0,1,2,…}, n q represents the starting frame of the q - th target.

[0011] S3. Dynamically select the review time frame and the review radar through a composite strategy, and verify the target by combining multi - domain resource scheduling within the dwell time;

[0012] S4. Calculate the overall miss - detection probability based on the review radar selected in step S3, and then construct a binary hypothesis test for multi - machine collaborative track deception to determine whether a re - review is required, thus completing the construction of the closed - loop detection mechanism.

[0013] Furthermore, the specific steps of step S3 are as follows:

[0014] S31. Selection of the review radar;

[0015] After the normal target tracking in step S2 and using the track of the fusion center to train the algorithm module, determine whether the current time frame is the -th frame. If not, continue with normal target tracking and use the track of the fusion center to train the algorithm module. If so, when the fusion center is at the -th frame, dynamically select the review radar m through a composite strategy, that is, select the review radar m by combining the UCB - 1 strategy and the ∈ - Greedy strategy among all radars. The expression is as follows:

[0016]

[0017] Among them, represents the obtained composite strategy, and ∈ represents the exploration rate in the ∈-Greedy strategy. represents the UCB-1 strategy value obtained according to the target distance and the verification times, and the calculation expression is as follows:

[0018]

[0019] Among them, (·) norm represents the min-max normalization of the maximum and minimum distances from all radars in the k-th frame to the q-th target, and C m,q represents the verification times of the m-th radar for the q-th target. The value of C m,q is constrained in the set {0.01, 1, 2,...}.

[0020] Then, at the frame, the q-th target is re-verified. If the selected verification radar m is not used to track the target q in the verification time frame , then the randomness of the composite strategy is used to select again until the selected verification radar m is used to track the target q in the verification time frame .

[0021] S32. Generation of multi-domain parameter scheduling strategy;

[0022] Use the predicted value of the fusion center to train the algorithm module U for the predicted position of the q-th target in the verification time frame m,q , and obtain the multi-domain resource scheduling strategy during the residence time of the verification radar m for the q-th target at the verification time frame , and verify the neighborhood of the predicted position of the target q in the corresponding verification time slot in the verification time frame;

[0023] Among them, the corresponding multi-domain parameter scheduling strategy is generated as follows:

[0024] First, model the multi-domain parameter scheduling as a Markov decision process including actions, states, and reward functions, and then use the Q-learning algorithm in reinforcement learning to solve the Markov decision process, and generate the multi-domain parameter scheduling strategy through the solution. The algorithm module for the m-th radar to track the q-th target is denoted as U m,q , and the action space in its Markov decision process has the following expression:

[0025]

[0026] Then, the optional action at the (n - 1)-th time step has the following expression:

[0027]

[0028] Among them, represents the set of optional powers of the radar, represents the set of optional pulse widths of the radar, represents the set of optional carrier frequencies of the radar, respectively represent the pulse width, frequency, and power of the pulse emitted by algorithm module U m,q at the nth time step, and for the kth frame, n = 1, 2,..., N k , N k represents the number of pulses in the pulse train of the kth frame.

[0029] Algorithm module U m,q The state space in the Markov decision process The expression is as follows:

[0030]

[0031] Among them, represents the first state after frequency hopping, represents the number of consecutive pulses in the same coherent processing interval since . The parameter K max represents the number of consecutive pulses in the same coherent processing interval since . The maximum duration step that the radar is allowed to stay on a single coherent processing interval after max . Within a short time interval of K times the pulse repetition interval, the jammer cannot quickly generate a new spoofing signal after verifying the multi-domain parameter agility of the radar. The networked radar can obtain an effective anti-spoofing window period. According to this constraint, when , the action of the algorithm module is set to a4, and the subsequent state transition is to The first pulse of the pulse train with n = 1 is also set to m,q (n), and the state at time step n is denoted as S

[0032] Algorithm module U m,q The reward function in the Markov decision process The expression is as follows:

[0033]

[0034] Among them, ω D represents the weight of the detection performance factor γ(n), ω T represents the weight of the task flexibility factor Ψ(n), and there is a constraint condition ω D +ω T = 1. The expression of the detection performance factor γ(n) in Equation (6) is as follows:

[0035]

[0036] Among them, represents the hopping cost. The expression of the task flexibility factor Ψ(n) in Equation (6) is as follows:

[0037]

[0038] Among them, PRF represents the pulse repetition frequency, represents the subarray aperture size, bw az and bw el represent the azimuth beam width and the elevation beam width respectively.

[0039] Then, the Q-learning algorithm in reinforcement learning is used to calculate the recursive optimized action selection strategy, and the expression is as follows:

[0040]

[0041] Among them, α represents the learning rate in Q-learning, and γ represents the discount rate in Q-learning.

[0042] Finally, the optimized action selection strategy is obtained through Equation (10), and multi-domain parameter scheduling is performed in the algorithm module U m,q , and the expression is as follows:

[0043]

[0044] Furthermore, the specific steps of step S4 are as follows:

[0045] First, estimate the signal-to-noise ratio of each coherent integration interval of the mth radar for tracking the qth target in the kth frame The expression is as follows:

[0046]

[0047] Among them, I w represents the number of pulses in the wth coherent integration interval, respectively represent the power, pulse width, and frequency of the ith pulse in this coherent integration interval, and are respectively selected from the radar optional power set the radar optional pulse width set Γ m , the radar optional carrier frequency set . G t represents the transmitting antenna gain, G r = G t represents the receiving antenna gain. c represents the speed of light, σ represents the non-fluctuating radar cross-sectional area of the target, k0T s represents the Boltzmann constant k0 and the system noise temperature T sThe product, represents the distance from the m-th radar to the q-th target at the k-th frame.

[0048] Based on Equation (12), the expression for calculating the signal-to-noise ratio of the echo of the transmitted pulse train for each coherent integration interval is as follows:

[0049]

[0050] where W represents the number of coherent integration intervals in a pulse train, and it has different values in different pulse trains.

[0051] Then, using the signal-to-noise ratio estimation in Equation (13), the corresponding detection probability P d (m, q, k) is obtained, and the expression is as follows:

[0052]

[0053] where P fa represents the false alarm probability of the radar. When the true target obtains a high cumulative signal-to-noise ratio in Equation (13), the detection probability P d (m, q, k) of the true target will be higher than that of the false target. Otherwise, during the review, the radar receiver only detects the noise power then the false target shows a lower cumulative signal-to-noise ratio.

[0054] The overall miss detection probability P miss (m, q, k) is estimated according to the result obtained from Equation (14), and the expression is as follows:

[0055]

[0056] where δ represents the length of the sliding window, which is used to set the number of frames involved in the calculation, and is set to δ = n q , n q represents the starting frame of the q-th target.

[0057] Furthermore, a binary hypothesis test for multi-aircraft cooperative track deception is constructed based on the overall miss detection probability obtained from Equation (15), and the expression is as follows:

[0058]

[0059] where V th represents a pre-set threshold, H T represents the hypothesis that the target is a true target, and H F represents the hypothesis that the target is a false target. When P miss (m, q, k) > V th , it is judged as a false target, and vice versa as a true target.

[0060] Based on the decision result of the binary hypothesis test, the verification radar feeds the false target back to the fusion center and marks it, returns to step S2 to reselect the verification time frame for re-verification, and resumes normal tracking of the identified true target at the next moment.

[0061] Then, according to the binary hypothesis test in Equation (16), the process of the networked radar against distributed multi-aircraft cooperative track deception interference is formed into an optimization problem and the parameters are optimized.

[0062] Establish an optimization problem to optimize the control decision parameters Make decisions on the pulse carrier frequency, pulse width, and power to minimize the expected time for identifying distributed multi-aircraft cooperative track deception interference. Then the expression of the optimization problem is as follows:

[0063]

[0064] Among them, E{·} represents taking the mathematical expectation of the expression in the brackets. The constraint conditions reflect the limitations of the radar system's detection performance, while for the constraint reflects the limitations brought by the radar mission requirements.

[0065] Finally, use the algorithm module obtained after steps S31 and S32 to indirectly and adaptively solve the optimization problem in Equation (17) under unknown scenarios.

[0066] The beneficial effects of the present invention: The method of the present invention first inserts a verification time slot during the conventional tracking process, and by transmitting a pulse train with multi-domain agility after reinforcement learning scheduling within the verification time slot, it can track the true target and identify the track deception of the false target in the presence of track deception. Then, through a composite strategy, it dynamically selects the verification time frame and the verification radar, combines the multi-domain resource scheduling within the dwell time to verify the target. Finally, based on the binary hypothesis test, it determines whether to re-verify, completing the construction of a closed-loop detection mechanism. The method of the present invention solves the problem that existing networked radars are difficult to identify such deception interference. The proposed verification mechanism can effectively identify distributed multi-aircraft cooperative track deception interference in complex multi-target scenarios, and the intelligent decision-making based on reinforcement learning further improves the system detection efficiency and resource utilization rate, and can be applied to fields such as target monitoring and radar countermeasure. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 is a flowchart of a method for a networked radar against distributed cooperative track deception interference based on reinforcement learning according to the present invention.

[0068] Figure 2 is a schematic diagram of a multi-target simulation scenario in an embodiment of the present invention.

[0069] Figure 3Schematic diagram for comparing transmitted pulses in the embodiments of the present invention.

[0070] Figure 4 Schematic diagram for selecting the review time frame and review radar in the multi-target simulation scenario in the embodiments of the present invention.

[0071] Figure 5 Schematic diagram of the cumulative miss detection probability P corresponding to Targets 1, 2, 3, and 4 in the embodiments of the present invention. miss of.

[0072] Figure 6 Schematic diagram for comparing the average detection probability obtained after the Monte Carlo experiment in the embodiments of the present invention.

[0073] Figure 7 Schematic diagram of the transmitted pulse parameters during the review process in the embodiments of the present invention. Detailed implementation manners

[0074] The method of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0075] To facilitate the description of the content of the method of the present invention, the following terms are first defined and explained:

[0076] Term 1: Review time frame

[0077] The review time frame refers to one or more time frames in the time frame of the networked radar used for the review process.

[0078] Term 2: Review radar

[0079] The review radar refers to the radar selected from the networked radar to perform the review process within the review time frame.

[0080] Term 3: Review mechanism

[0081] The review mechanism refers to the overall mechanism composed of all the steps mentioned in the present invention.

[0082] Term 4: Track planning

[0083] Track planning refers to the multi-jammer cooperative planning carried out during the distributed multi-aircraft cooperative track deception jamming process. Due to the characteristics of a single aircraft performing deception jamming on the main lobe of a single radar, in order to increase the number of false targets created and make full use of the existing jammers for deployment. It is mainly realized through the design of tracks and related planning control algorithms.

[0084] As Figure 1 shown, the flowchart of a method for a networked radar to resist distributed cooperative track deception jamming based on reinforcement learning according to the present invention is as follows:

[0085] S1. In the scenario of distributed multi-aircraft collaborative deception jamming, obtain the parameter settings of the networked radar system at a certain moment, and initialize the algorithm module according to the training parameters and the parameter settings of the networked radar system;

[0086] S2. The target forms a track at the fusion center, and a review time slot is inserted during the tracking process. The fusion center allocates resources to track the target;

[0087] For the target that forms a track, a review time slot is inserted during the tracking process, that is, for any target q that forms a track at the networked radar fusion center, after the fusion center allocates resources to track the target, a review time frame is dynamically selected as the inserted review time slot, that is, according to the pre-set review interval n c , the review time frame of the q-th target is randomly selected within [n q +b·n c +1, n q +(b + 1)·n c . Before the review time frame , the normal target tracking is carried out and the algorithm module corresponding to the q-th target is trained using the track at the fusion center. The algorithm module is denoted as U m,q , m = 1,... M.

[0088] where M represents the number of radar points in the radar network. b = {0, 1, 2,...}, and n q represents the starting frame of the q-th target.

[0089] S3. Dynamically select the review time frame and review radar through a composite strategy, and verify the target by combining multi-domain resource scheduling within the dwell time;

[0090] S4. Calculate the overall miss detection probability based on the review radar selected in step S3, and then construct a binary hypothesis test for multi-aircraft collaborative track deception to determine whether a re-review is required, thus completing the construction of the closed-loop detection mechanism.

[0091] In this embodiment, the specific steps of step S1 are as follows:

[0092] In this embodiment, according to the target initial state parameters in Table 1, a multi-target simulation scenario of multi-aircraft collaborative deception jamming is constructed as Figure 2 shown, and the time interval of each frame is T = 2s.

[0093] Table 1

[0094]

[0095] Figure 2Each track in the shown scenario can be replaced with a target of any motion model, and the number of targets can be increased or decreased. The corresponding parameters of the networking radar system are shown in Table 2, and the training parameter settings are shown in Table 3.

[0096] Table 2

[0097]

[0098] Table 3

[0099]

[0100] The multi-domain agile pulse train emitted during the review process of this embodiment is as Figure 3 shown. In the figure, f i and τ i respectively represent the radar pulse carrier frequency and pulse width. W is the number of coherent accumulation intervals in a pulse train in Equation (13), and has different values in different pulse trains.

[0101] In this embodiment, the specific steps of step S3 are as follows:

[0102] S31. Review radar selection;

[0103] After the normal target tracking in step S2 and using the track training algorithm module of the fusion center, it is judged whether the current time frame is the frame. If not, continue with normal target tracking and use the track training algorithm module of the fusion center. If so, when the fusion center is at the frame, the review radar m is dynamically selected through a composite strategy, that is, the review radar m is selected by combining the UCB-1 strategy and the ∈-Greedy strategy among all radars. The expression is as follows:

[0104]

[0105] Among them, represents the obtained composite strategy, ∈ represents the exploration rate in the ∈-Greedy strategy, represents the UCB-1 strategy value obtained according to the target distance and verification times. The calculation expression is as follows:

[0106]

[0107] Among them, (·) norm represents the min-max normalization of the maximum and minimum distances from all radars to the qth target in the kth frame. C m,q represents the verification times of the mth radar for the qth target. To avoid regular divergence, the value of C m,q is constrained in the set {0.01, 1, 2,...}.

[0108] Then at the The frame rechecks the q-th target. If the selected recheck radar m is not used to track target q in the recheck time frame it is selected again using the randomness of the composite strategy until the selected recheck radar m is used to track target q in the recheck time frame .

[0109] S32. Generation of multi-domain parameter scheduling strategy;

[0110] Use the predicted value of the fusion center to train the algorithm module U for the predicted position of the q-th target in the recheck time frame to obtain the multi-domain resource scheduling strategy during the dwell time of the recheck radar m on the q-th target at the recheck time frame, and recheck the neighborhood of the predicted position of target q in the corresponding recheck time slot in the recheck time frame; m,q Among them, the corresponding multi-domain parameter scheduling strategy is generated as follows:

[0111]

[0112] First, model the multi-domain parameter scheduling as a Markov decision process including actions, states, and reward functions, and then use the Q-learning algorithm in reinforcement learning to solve the Markov decision process to generate the multi-domain parameter scheduling strategy. The algorithm module for the m-th radar to track the q-th target is denoted as U m,q , and the action space in its Markov decision process is expressed as follows:

[0113]

[0114] Then the optional actions at the (n - 1)-th time step (pulse) are expressed as follows:

[0115]

[0116] Among them, represents the set of optional powers of the radar, represents the set of optional pulse widths of the radar, represents the set of optional carrier frequencies of the radar, respectively represent the pulse width, frequency, and power of the pulse emitted by the algorithm module U m,q at the n-th time step, and for the k-th frame, n = 1, 2,..., N k , N k represents the number of pulses in the pulse train of the k-th frame.

[0117] The state space m,q of the algorithm module U in the Markov decision process is expressed as follows:

[0118] ​​​

[0119] Among them, represents the first state after frequency hopping, represents since the number of consecutive pulses in the same coherent processing interval ( pulses). The parameter K max represents the maximum duration step that the radar is allowed to stay on a single coherent processing interval since . In a short time interval of K max times the pulse repetition interval, since the jammer cannot quickly generate a new deception signal after verifying the multi-domain parameter agility of the radar, the networked radar can obtain an effective anti-deception window period. According to this constraint, when , the action setting of the algorithm module is a4, and the subsequent state transition is The first pulse of the pulse train with n = 1 is also set to the state of. The state at time step n is denoted as S m,q (n), and

[0120] The algorithm module U m,q The reward function in the Markov decision process is expressed as follows:

[0121]

[0122] Among them, ω D represents the weight of the detection performance factor γ(n), ω T represents the weight of the task flexibility factor Ψ(n), and there is a constraint condition ω D +ω T = 1. The expression of the detection performance factor γ(n) in Equation (6) is as follows:

[0123]

[0124] Among them, represents the frequency hopping cost. The expression of the task flexibility factor Ψ(n) in Equation (6) is as follows:

[0125]

[0126] Among them, PRF represents the pulse repetition frequency, represents the subarray aperture size, bw az and bw el represent the azimuth beam width and elevation beam width respectively.

[0127] Then, the Q-learning algorithm in reinforcement learning is used to calculate the recursive optimization action selection strategy, and the expression is as follows:

[0128]

[0129] Among them, α represents the learning rate in Q-learning, and γ represents the discount rate in Q-learning.

[0130] Finally, the optimized action selection strategy is obtained through Equation (10), and multi-domain parameter scheduling is performed in algorithm module U m,q as follows:

[0131]

[0132] In this embodiment, the specific steps of step S4 are as follows:

[0133] First, estimate the signal-to-noise ratio of each coherent integration interval of the m-th radar for tracking the q-th target in the k-th frame The expression is as follows:

[0134]

[0135] Among them, I w represents the number of pulses in the w-th coherent integration interval, respectively represent the power, pulse width, and frequency of the i-th pulse in this coherent integration interval, and are respectively selected from the radar selectable power set the radar selectable pulse width set Γ m , the radar selectable carrier frequency set . G t represents the transmitting antenna gain, G r = G t represents the receiving antenna gain. c represents the speed of light, σ represents the non-fluctuating radar cross section of the target, k0T s represents the product of the Boltzmann constant k0 and the system noise temperature T s , represents the distance from the m-th radar to the q-th target at the k-th frame.

[0136] Then, based on Equation (12), the expression for calculating the echo signal-to-noise ratio of the transmitted pulse train for the signal-to-noise ratio of each coherent integration interval is as follows:

[0137]

[0138] Among them, W represents the number of coherent integration intervals in a pulse train, and has different values in different pulse trains.

[0139] Then, using the signal-to-noise ratio estimation in Equation (13), the corresponding detection probability P d (m, q, k) is obtained, and the expression is as follows:

[0140]

[0141] Among them, P fa represents the false alarm probability of the radar. When the true target obtains a high cumulative signal-to-noise ratio in Equation (13), the detection probability P d (m, q, k) of the true target will be higher than that of the false target. Otherwise, during the review, the radar receiver only detects the noise power then the false target shows a lower cumulative signal-to-noise ratio.

[0142] Estimate the overall missed detection probability P miss (m, q, k) according to the result obtained from Equation (14), and the expression is as follows:

[0143]

[0144] Among them, δ represents the length of the sliding window, which is used to set the number of frames participating in the calculation, and is set to δ = n q , n q represents the starting frame of the q-th target.

[0145] Then, construct a binary hypothesis test for multi-aircraft cooperative track deception based on the overall missed detection probability obtained from Equation (15), and the expression is as follows:

[0146]

[0147] Among them, V th represents a pre-set threshold, H T represents the hypothesis that this target is a true target, and H F represents the hypothesis that this target is a false target. When P miss (m, q, k) > V th , it is judged as a false target, otherwise it is judged as a true target.

[0148] Based on the decision result of the binary hypothesis test, the review radar feeds the false target back to the fusion center and marks it, returns to step S2 to re-select the review time frame for re-review, and resumes normal tracking for the identified true target at the next moment.

[0149] Then, according to the binary hypothesis test in Equation (16), form an optimization problem for the process of the networked radar against distributed multi-aircraft cooperative track deception interference and perform parameter optimization.

[0150] Establish an optimization problem to optimize the control decision parameters Make decisions on the pulse carrier frequency, pulse width, and power to minimize the expected time for identifying distributed multi-aircraft cooperative track deception interference. Then, the expression of the optimization problem is as follows:

[0151]

[0152] Among them, E{·} represents taking the mathematical expectation of the expression in the brackets. The constraint conditions reflect the limitations of the radar system's detection performance, while for the constraint it reflects the limitations brought by the radar mission requirements.

[0153] Since the objective function does not have an analytical form, it is impossible to obtain an explicit solution. Finally, the algorithm module obtained after steps S31 and S32 (the algorithm module based on reinforcement learning) is used to indirectly and adaptively solve the optimization problem in Equation (17) in an unknown scenario.

[0154] In this embodiment, the verification of the method of the present invention is completed in a multi-target scenario with distributed multi-aircraft cooperative track deception interference, and the results are as Figures 4 to 7 shown.

[0155] Figure 4 The comparison in a single simulation experiment shows that the reinforcement learning-based solution of the method of the present invention is superior to the random method in accurate and efficient target recognition and can correctly distinguish true and false targets. Figure 4 (a) is the method of the present invention, Figure 4 (b) is the random method. It can be seen that the reinforcement learning-based method can correctly identify the true target 1 with only one review in the 3rd frame, while the random method requires an additional review process and cannot complete the identification until the end of the 8th frame. Similarly, for the true target 3, the reinforcement learning-based method can identify the target as true before the 19th frame with only 4 review processes, while the random method cannot determine that the target is true until the 24th frame. Figure 4 As

[0156] shown, the cumulative miss detection probabilities P Figure 5 corresponding to targets 1, 2, 3, and 4 are respectively as miss shown in Figure 5 (a)(b)(c)(d). Among them, the true targets have high signal-to-noise ratio echo pulses and quickly approach the threshold V th as Figure 5 (a)(c)(d) shown. Different from the true targets, the false targets (target 2 and target 5) because of similar Figure 5 P shown in miss do not meet the P miss < V th criterion and are regularly verified.

[0157] In this embodiment, Figure 4 it also shows the advantage of the reinforcement learning-based solution in the present invention method in terms of resource scheduling efficiency. As Figure 5 (d) shown, even for the true target 4 with the same review frames and review radars, the random method still shows a higher P miss, thus making it easier to misclassify true targets.

[0158] As Figure 6 shown, in addition to improving the efficiency of the review mechanism, the reinforcement learning-based solution of the method of the present invention also enhances the radar detection performance and achieves a continuously stable high detection probability. To prove this, the detection probabilities for a target far away ( Figure 6 (a) Target 1 approaching Radar 1) and a target approaching ( Figure 6 (b) Target 3 approaching Radar 1) are given, and the results are taken as the average of 100 Monte Carlo simulations. In both cases, the reinforcement learning-based solution is always superior to the random method, achieving a higher detection probability and better tracking ability for true targets.

[0159] Figure 7 Then, the source of the above advantages is demonstrated from the aspect of pulse parameters by the power and carrier frequency of the pulse (for convenience of display, the carrier frequencies not selected by the radar are not drawn in the figure). Figure 7 (a) Schematic diagram of the transmitted pulse parameters (pulse carrier frequency and power) during the review process, Figure 7 (b) Schematic diagram of the transmitted pulse parameters (pulse carrier frequency and pulse width) during the review process. It can be seen from Figure 7 that the output strategy of the algorithm module expands the hopping frequency interval within the maximum state index K max range, optimizes the detection performance according to the signal-to-noise ratio, and at the same time weighs the power consumption to maintain task flexibility. These parameter adjustments enable the radar system to adapt to changing conditions and enhance the target detection ability.

[0160] In summary, the method of the present invention models the adaptive decision-making anti-jamming process of the networked radar in an unknown scenario as a Markov decision process, and through the reinforcement learning algorithm and strategy, realizes the dynamic optimization and sequential decision-making of the radar, timing, and pulse parameters, thereby forming a closed-loop anti-jamming mechanism, which can effectively identify distributed multi-aircraft cooperative track deception jamming. The method of the present invention can effectively identify distributed multi-aircraft cooperative track deception jamming in a complex multi-target scenario, and the intelligent sequential decision-making based on reinforcement learning can improve the system detection efficiency and resource utilization rate, providing a new technical path (an expandable technical path) for the networked radar to resist distributed cooperative deception, and can be used to cope with distributed multi-aircraft cooperative track deception jamming generated by different multi-aircraft cooperation and track planning methods.

[0161] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A method for a networking radar to resist distributed cooperative track deception jamming based on reinforcement learning, the specific steps are as follows: S1. In the scenario of distributed multi-aircraft cooperative deception jamming, obtain the parameter settings of the networking radar system at a certain moment, and initialize the algorithm module according to the training parameters and the parameter settings of the networking radar system; S2. The target forms a track at the fusion center, and a review time slot is inserted during the tracking process. The fusion center allocates resources to track the target; For a target forming a track, a review time slot is inserted during the tracking process. That is, for any target q that forms a track at the networked radar fusion center, after the fusion center allocates resources to track the target, it dynamically selects a review time frame. As the inserted review time slot, that is, according to the preset review interval n c , the review time frame of the q-th target is randomly selected between [n q +b·n c +1, n q +(b + 1)·n c ; before the review time frame , normal target tracking is performed and the algorithm module corresponding to the q-th target is trained using the track at the fusion center. The algorithm module is denoted as U m,q , m = 1, … M; where M represents the number of radar points in the radar network; b = {0, 1, 2, …}, n q represents the starting frame of the q-th target; S3. Dynamically select the review time frame and review radar through a composite strategy, and verify the target by combining multi-domain resource scheduling within the dwell time; S4. Calculate the overall miss detection probability based on the review radar selected in step S3, and then construct a binary hypothesis test for multi-aircraft cooperative track deception to determine whether a re-review is required, and complete the construction of the closed-loop detection mechanism.

2. A method for a networking radar to resist distributed cooperative track deception interference based on reinforcement learning according to claim 1, characterized in that, The specific steps of step S3 are as follows: S31. Review radar selection; After the normal target tracking in step S2 and using the track training algorithm module of the fusion center, it is judged whether the current time frame is the frame. If not, continue the normal target tracking and use the track training algorithm module of the fusion center. If so, at the frame, the fusion center dynamically selects the verification radar m through a composite strategy, that is, combines the UCB-1 strategy and the ∈-Greedy strategy to select the verification radar m among all radars. The expression is as follows: Among them, represents the obtained composite strategy, and ∈ represents the exploration rate in the ∈-Greedy strategy. represents the UCB-1 strategy value obtained according to the target distance and the number of verification times, and the calculation formula is as follows: Among them, (·) norm represents the min-max normalization of the maximum and minimum distances from all radars in the k-th frame to the q-th target, C m,q represents the verification times of the m-th radar for the q-th target; C m,q is constrained to take values in the set {0.01, 1, 2,...}; Then, at the frame, recheck the q-th target. If the selected recheck radar m is not used to track target q in the recheck time frame , then use the randomness of the composite strategy to select again until the selected recheck radar m is used to track target q in the recheck time frame ; S32. Generation of multi-domain parameter scheduling strategy; Use the predicted value of the fusion center to train the algorithm module U for the predicted position of the q-th target in the review time frame to obtain the multi-domain resource scheduling strategy of the review radar m during the dwell time of the q-th target in the review time frame m,q and perform a review on the neighborhood of the predicted position of target q in the corresponding review time slot in the review time frame; ​ Among them, to generate the corresponding multi-domain parameter scheduling strategy, the specific steps are as follows: First, model the multi-domain parameter scheduling as a Markov decision process that includes actions, states, and a reward function, and then use the Q-learning algorithm in reinforcement learning to solve the Markov decision process. A multi-domain parameter scheduling policy is generated through the solution. The algorithm module for the m-th radar to track the q-th target is denoted as U m,q , and the action space in its Markov decision process The expression is as follows: The actions available at the (n-1)-th time step The expression is as follows: Among them, represents the set of optional radar powers, represents the set of optional radar pulse widths, represents the set of optional radar carrier frequencies, respectively represent the pulse width, frequency, and power of the pulse emitted by algorithm module U m,q at the nth time step, and for the kth frame, n = 1, 2,..., N k , N k represents the number of pulses in the kth frame pulse train; Algorithm module U m,q State space in Markov decision process The expression is as follows: Among them, represents the first state after frequency hopping, represents the number of consecutive pulses in the same coherent processing interval since ; the parameter K max represents the maximum duration step that the radar is allowed to stay on a single coherent processing interval since ; within a short time interval of K max times the pulse repetition interval, the jammer cannot quickly generate a new spoofing signal after rechecking the multi-domain parameter agility of the radar, and the networked radar can obtain an effective anti-spoofing window period. According to this constraint, when , the action setting of the algorithm module is a4, and the subsequent state transition is to The first pulse of the pulse train with n = 1 is also set to The state at time step n is denoted as S m,q (n), and Algorithm module U m,q Reward function in the Markov decision process The expression is as follows: Among them, ω D represents the weight of the detection performance factor γ(n), ω T represents the weight of the task flexibility factor Ψ(n), and there is a constraint condition ω D +ω T = 1; The expression of the detection performance factor γ(n) in Equation (6) is as follows: Among them, represents the hopping cost; the expression of the task flexibility factor Ψ(n) in Equation (6) is as follows: Among them, PRF represents the pulse repetition frequency, represents the subarray aperture size, bw az and bw el represent the azimuth beamwidth and the elevation beamwidth respectively; Then use the Q-learning algorithm in reinforcement learning to calculate the recursive optimization action selection strategy, and the expression is as follows: Among them, α represents the learning rate in Q-learning, and γ represents the discount rate in Q-learning; Finally, the optimized action selection strategy is obtained through Equation (10) and multi-domain parameter scheduling is performed in algorithm module U m,q The expression is as follows:

3. A method for anti-distributed cooperative track spoofing interference of a networking radar based on reinforcement learning according to claim 1, characterized in that, The specific steps of step S4 are as follows: First, estimate the signal-to-noise ratio of each coherent integration interval of the m-th radar for tracking the q-th target in the k-th frame. The expression is as follows: Among them, I w represents the number of pulses in the w-th coherent integration interval, respectively represent the power, pulse width, and frequency of the i-th pulse in this coherent integration interval, and are respectively selected from the radar selectable power set the radar selectable pulse width set Γ m , the radar selectable carrier frequency set ; G t represents the transmitting antenna gain, G r = G t represents the receiving antenna gain; c represents the speed of light, σ represents the non-fluctuating radar cross-section of the target, k0T s represents the product of the Boltzmann constant k0 and the system noise temperature T s ; represents the distance from the m-th radar to the q-th target at the k-th frame; Then based on Equation (12), the echo signal-to-noise ratio of the transmitted pulse train is calculated for each coherent integration interval, and the expression is as follows: Among them, W represents the number of coherent integration intervals in a pulse train, and has different values in different pulse trains; Then, using the signal-to-noise ratio estimation in Equation (13), the corresponding detection probability P d (m, q, k) is obtained, and the expression is as follows: Among them, P fa represents the false alarm probability of the radar; when the true target obtains a high cumulative signal-to-noise ratio in Equation (13), the detection probability P d (m, q, k) of the true target will be higher than that of the false target. Otherwise, during the review, the radar receiver only detects the noise power then the false target shows a lower cumulative signal-to-noise ratio; Estimate the overall missed detection probability \(P\) according to the result obtained from Equation (14) miss (m, q, k), and the expression is as follows: where δ represents the length of the sliding window, which is used to set the number of frames involved in the calculation, and is set to δ = n q , n q represents the starting frame of the q-th target; Then construct a binary hypothesis test for multi-aircraft cooperative track deception based on the overall miss detection probability obtained from Equation (15), and the expression is as follows: where, V th represents a preset threshold value, H T represents the assumption that the target is a true target, H F represents the assumption that the target is a false target; when P miss (m, q, k) > V th it is determined as a false target, otherwise it is determined as a true target; Based on the binary hypothesis test decision result, the review radar feeds the false target back to the fusion center and marks it, returns to step S2 to re-select the review time frame for re-review, and resumes normal tracking for the target identified as a true target at the next moment; Then according to the binary hypothesis test in Equation (16), form an optimization problem for the process of the networking radar to resist distributed multi-aircraft cooperative track deception jamming and perform parameter optimization; Establish an optimization problem to optimize the control decision parameters Make decisions on the pulse carrier frequency, pulse width, and power to minimize the expected time for distributed multi-aircraft cooperative track deception interference discrimination. Then, the expression of the optimization problem is as follows: Among them, E{·} represents taking the mathematical expectation of the expression in the brackets. The constraint condition reflects the limitations of the radar system's detection performance, and for the constraint reflects the limitations brought by the radar mission requirements; Finally, use the algorithm module obtained after steps S31 and S32 to indirectly and adaptively solve the optimization problem in Equation (17) in an unknown scenario.