A radar multi-dimensional resource management method for multi-target tracking in a dynamic electromagnetic environment
By acquiring jammer information from the radar and combining it with target state prediction, the transmission frequency and power resources are decomposed and optimized, solving the problem of inaccurate radar resource management in dynamic electromagnetic environments and improving target tracking performance and survivability.
Patent Information
- Application Number
- CN202410466168.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-18
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-04-18
AI Technical Summary
Existing radar resource management models fail to effectively consider the impact of interference on target tracking in dynamic electromagnetic environments, resulting in inaccurate resource allocation and weakening the radar's survivability and target tracking performance.
By acquiring information about the jammer and combining it with target state prediction and active anti-jamming MDP model, the radar MRM scheme is decomposed into sub-problems of transmission frequency and transmission power. Reinforcement learning method is used to optimize resource allocation, forming a comprehensive radar multi-dimensional resource management method.
In dynamic electromagnetic environments, the radar's target tracking performance and survivability are improved, its sensitivity to interference is reduced, and the overall design is more in line with practical application requirements.
Smart Images

Figure CN118519137B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of radar detection, and particularly relates to a radar multi-dimensional resource management (MRM) method for multi-target tracking in a dynamic electromagnetic environment. BACKGROUND
[0002] In recent years, due to limited radar system resources, resource management is crucial to improve the performance of MTT (Multiple Target Tracking). When multiple targets are tracked, the resource management of MTT can find a compromise between target tracking performance and limited system resources. Technically, this problem can be described as an optimization model to enhance resource use efficiency.
[0003] Currently, there have been some researches on various aspects of the above problem, and resource management schemes have been designed to improve the performance of MTT. For example, some researchers have proposed a joint beam and power allocation solution to improve the performance of MTT, and developed a robust power allocation algorithm to meet the expected MTT performance when the available power is insufficient; some researchers have constructed a dual-objective constrained optimization framework for integrated search and tracking tasks application, so that the radar can obtain better target search and tracking task utility by solving the optimization problem; some researchers have designed a scheme to reduce the interception probability while meeting the predetermined target tracking performance, which plays an important role in improving the survivability of the radar.
[0004] Current radar resource management schemes are built in a clean electromagnetic environment and only utilize prior information about targets to build models. With the continuous development of modern electronic countermeasures (ECM) and electronic counter-countermeasures (ECCM), jamming has become a major means of electronic warfare, usually aiming to mask the useful information of targets or deceive the radar system to greatly reduce its performance. When multiple targets with jammers appear in the monitoring area, the radar needs to coordinate its transmission resources to counter the jammers or improve the MTT performance. The existing radar resource management model neither considers the impact of jamming on target tracking nor considers the means to counter the jamming, so the resource allocation solution generated thereby is inaccurate and the survivability of the radar equipment is greatly weakened. Therefore, it is crucial to build a corresponding resource management model for the jamming environment. In reality, the jamming strategy of the jammer is usually unknown and dynamic to the radar. The radar has no prior information about the jammer, which makes it more difficult to implement resource management in a jamming environment. Therefore, most current work focuses more on dealing with unknown and dynamic jamming environments. Among them, some researchers have proposed an active countermeasure method based on reinforcement learning (RL) to adjust the transmission resources of the radar in advance to avoid being jammed, in which frequency agility (FA) is an effective method; some researchers have designed an anti-jamming model based on RL to adjust the transmission frequency to counter intelligent jammers and improve the detection probability; for moving targets, some researchers have adopted a similar framework to manage frequency resources to avoid being interfered by a communication system using the same frequency band while providing a higher signal to interference plus noise ratio (SINR) and higher bandwidth utilization.
[0005] The above research work shows that resource management and anti-jamming can directly or indirectly improve target tracking performance. However, these existing works either perform resource management only in a clean and static environment or counter jamming only in a dynamic environment, without fully combining and utilizing both and making the MTT task obtain the desired performance. SUMMARY
[0006] In order to solve the above problems existing in the prior art, the present application provides a radar multi-dimensional resource management method for multi-target tracking in a dynamic electromagnetic environment.
[0007] The technical problem to be solved by the present application is solved by the following technical scheme:
[0008] The application provides a radar multi-dimensional resource management method for multi-target tracking in a dynamic electromagnetic environment, comprising the following steps:
[0009] Step 1: obtaining jammer information of a jammer, wherein the jammer information comprises frequency and jamming mode of the jammer;
[0010] Step 2: obtaining predicted state information of a qth target in a kth continuous time period based on state estimation information of the qth target in a (k-1) th continuous time period and a target state transition matrix, wherein 1≤q≤Q, and Q represents the total number of targets;
[0011] Step 3: introducing the jammer information into an active anti-jamming MDP model, and combining the predicted state information of the qth target in the kth continuous time period to obtain a radar MRM scheme model to be solved in the kth continuous time period;
[0012] Step 4: decomposing the radar MRM scheme model to be solved in the kth continuous time period into a first sub-problem related to a transmission frequency and a second sub-problem related to a transmission power, and obtaining an MRM scheme according to optimal transmission schemes and random transmission schemes determined based on the first sub-problem and the second sub-problem;
[0013] Step 5: determining state estimation information of the qth target in the kth continuous time period based on observation data obtained by a radar executing the MRM scheme;
[0014] Step 6: returning to step 1, entering a (k+1) th continuous time period, and continuing until a requirement of stopping operation is reached.
[0015] Optionally, the predicted state information of the qth target in the kth continuous time period is represented as:
[0016] ξ q,k|k-1 =Fξ q,k-1|k-1
[0017] wherein ξ q,k|k-1 represents the predicted state information of the qth target in the kth continuous time period, ξ q,k-1|k-1 represents state estimation information of the qth target in the (k-1) th continuous time period, and F represents a target state transition matrix, wherein for an NCV model, F is denoted as F NCV , for an NCT model, F is denoted as F NCT , the target state transition matrix F NCV and the target state transition matrix F NCT are respectively represented as:
[0018]
[0019]
[0020] wherein I2 represents a 2-order unit matrix, denotes the Kronecker operator, T denotes the time interval between consecutive time instants, and w denotes the rotational angular velocity.
[0021] Optionally, the step 3 comprises:
[0022] At the kth consecutive time instant, the action taken by the radar on the qth target is obtained, denoted as:
[0023]
[0024] wherein, denotes the action taken by the radar on the qth target at the kth consecutive time instant, denotes the carrier frequency of the signal transmitted by the radar on the qth target at the kth consecutive time instant when taking the action , fi denotes the minimum carrier frequency selectable by the radar, and Df denotes the step size between two consecutive carrier frequencies, Df > 0, denotes the action space of the radar, denotes the number of carrier frequencies selectable by the radar;
[0025] At the kth consecutive time instant, the action space of the jammer on the qth target is obtained, denoted as:
[0026]
[0027] wherein, denotes the action taken by the jammer on the qth target at the kth consecutive time instant, denotes the number of available jamming actions, is the action space of the jammer, at the kth consecutive time instant, if it indicates that the radar is not jammed, if it indicates that the jammer releases a frequency-agile jamming signal with the frequency , if it indicates that the jammer releases a barrage jamming signal, is the jamming signal frequency transmitted by the jammer on the qth target at the kth consecutive time instant when taking the action ;
[0028] According to the action taken by the radar on the qth target and the action space of the jammer on the qth target, the interaction information of the radar and the jammer on the qth target at the kth consecutive time instant is obtained, denoted as:
[0029]
[0030] wherein, represents the mutual information of the radar and the jammer on the qth target at the kth continuous time instant;
[0031] At the kth continuous time instant, based on the last M groups of observation values in the mutual information of the radar and the jammer on the qth target at the kth continuous time instant, the state of the radar for the qth target at the kth continuous time instant is obtained, and the state of the radar for the qth target at the kth continuous time instant is represented as:
[0032]
[0033]
[0034] wherein, represents the state of the radar for the qth target at the kth continuous time instant, S is a state space of the radar, represents an action taken by the qth target at the k-Mth continuous time instant, represents an action taken by the jammer on the qth target at the k-Mth continuous time instant, represents an action taken by the qth target at the k-1th continuous time instant, represents an action taken by the jammer on the qth target at the k-1th continuous time instant;
[0035] After the radar takes the action for the qth target, based on the maximum likelihood method, the final reward function is obtained according to the sum of all rewards obtained by the radar from the state of the radar for the qth target at the kth continuous time instant to the state of the radar for the qth target at the k+1th continuous time instant and the number of transitions of the radar from the state of the radar for the qth target at the kth continuous time instant to the state of the radar for the qth target at the k+1th continuous time instant;
[0036] Based on the maximum likelihood method, the state transition probability is obtained according to the number of transitions of the radar from the state of the radar for the qth target at the kth continuous time instant to the state of the radar for the qth target at the k+1th continuous time instant, and the state transition probability is represented as:
[0037]
[0038] wherein, represents the state transition probability, and satisfies if then
[0039] The Bayesian information matrix of the kth continuous moment is predicted according to the predicted state information of the qth target of the kth continuous moment, the final reward function and the state transition probability, so as to determine the radar MRM scheme model to be solved at the kth continuous moment according to the predicted Bayesian information matrix of the kth continuous moment.
[0040] Optionally, the final reward function is represented as:
[0041]
[0042]
[0043] wherein, represents the final reward function, represents the sum of all rewards obtained by the radar from the state of the qth target at the kth continuous moment to the state of the qth target at the k+1th continuous moment, represents the number of transitions from the state of the qth target at the kth continuous moment to the state of the qth target at the k+1th continuous moment, represents the state of the qth target at the kth continuous moment, and if reward function at the bth calculation reward function is represented as:
[0044]
[0045]
[0046] wherein, D q,k|k represents the distance estimation value between the qth target and the radar at the kth continuous moment, represents the transmission power of the radar for the qth target at the kth moment, N R represents the number of pulses in one coherent processing interval, G R,t represents the gain of the transmission antenna, represents the radar scattering cross section area of the qth target at the carrier frequency under, l R,t represents the signal wavelength, D q,k is the distance between the qth target and the radar, represents the power density of the interference signal, G J,t represents the transmission antenna gain of the jammer, l J,t represents the wavelength of the jammer, B R,r represents the bandwidth of the radar receiving signal, G R,r represents the gain of the receiving antenna, GR,p represents a processing gain, k0represents a Boltzmann constant, T N represents a noise temperature of the radar, F R,r represents a noise figure.
[0047] Optionally, a Bayesian information matrix at the kth continuous moment is predicted according to the predicted state information of the qth target at the kth continuous moment, the final reward function and the state transition probability, so as to determine a radar MRM scheme model to be solved at the kth continuous moment according to the predicted Bayesian information matrix at the kth continuous moment, and the radar MRM scheme model to be solved at the kth continuous moment is represented as:
[0048] a Bayesian information matrix at the kth continuous moment is predicted according to the predicted state information of the qth target at the kth continuous moment, the final reward function and the state transition probability;
[0049] a predicted BCRLB is obtained according to an inverse of the Bayesian information matrix at the kth continuous moment;
[0050] an MTT index is obtained according to the predicted BCRLB, and the MTT index is represented as:
[0051]
[0052] wherein, represents an MTT index, Trrepresents a trace of a matrix, represents a predicted BCRLB;
[0053] a target function of a radar MRM scheme model is obtained according to the MTT index, and the target function of the radar MRM scheme model is represented as:
[0054]
[0055] wherein, represents a target function of a radar MRM scheme model, P R,k represents a set of transmission powers of the radar for all targets at the kth continuous moment, a k represents a set of actions of the radar for all targets at the kth continuous moment;
[0056] the radar MRM scheme model to be solved at the kth continuous moment is determined according to the target function of the radar MRM scheme model, and the radar MRM scheme model to be solved at the kth continuous moment is represented as:
[0057]
[0058]
[0059]
[0060] wherein, P Total denotes the total amount of power resources that the radar can allocate, P min denotes the lower limit of power resources that the radar can allocate for a single target, P max denotes the upper limit of power resources that the radar can allocate for a single target.
[0061] Optionally, the Bayesian information matrix at the kth consecutive time instant is denoted as:
[0062]
[0063] wherein, denotes the Bayesian information matrix at the kth consecutive time instant, Q q,k-1 denotes the covariance of process noise at the kth consecutive time instant, denotes the 2nd-order identity matrix, denotes the Kronecker operator, T denotes the time interval between consecutive time instants, q0 denotes the process noise intensity, F denotes the target state transition matrix, denotes the Bayesian information matrix at the (k-1)th consecutive time instant, denotes the transmit power of the radar for the qth target at the (k-1)th time instant, denotes the action taken by the radar for the qth target at the (k-1)th consecutive time instant, T denotes the transpose of a matrix or a vector, q,k denotes the Jacobian matrix of the radar for the qth target at the kth consecutive time instant, q.k denotes the covariance matrix of observation noise of the radar for the qth target at the kth consecutive time instant, denotes the expected return, denotes the final reward function, q,kk-1 denotes the predicted distance between the qth target and the radar at the kth consecutive time instant, R,r denotes the bandwidth of the received signal of the radar, R,NN denotes the 3dB receiving beam width, R,d denotes the dwell time.
[0064] Optionally, the step 4 comprises:
[0065] decomposing the radar MRM scheme model to be solved at the kth consecutive time instant into a first sub-problem related to the transmit frequency and a second sub-problem related to the transmit power, the first sub-problem and the second sub-problem being respectively denoted as:
[0066]
[0067]
[0068] where, denotes the first sub-problem, denotes the second sub-problem, * denotes the optimal solution obtained by corresponding parameters;
[0069] RL-based policy iteration method, according to the new policy obtained by updating the value function of the policy constructed by the Bellman equation, the optimal solution of the first sub-problem is obtained;
[0070] gradient projection method, the optimal solution of the second sub-problem is obtained;
[0071] According to the optimal transmission scheme and the random transmission scheme determined by the first sub-problem and the second sub-problem, the MRM scheme is obtained, which is represented as:
[0072]
[0073] where, denotes the random transmission scheme, denotes a set of transmission frequencies selected from a uniform distribution, denotes a set of transmission powers selected from a uniform distribution of total resources, denotes the optimal transmission scheme, denotes a set of optimal solutions of the first sub-problem of all targets, denotes the optimal solution of the second sub-problem, e is the greed rate, 0£e£1.
[0074] Optionally, the step 5 comprises:
[0075] The radar executes the MRM scheme to obtain the observation data of the radar on the qth target at the kth continuous time;
[0076] According to the measurement function of the observation data of the radar on the qth target at the kth continuous time and the predicted state information of the qth target at the kth continuous time, the innovation between the observation data and the predicted data is obtained, which is represented as:
[0077] v q,k =z q,k -h(ξ q,k|k-1 )
[0078] where, z q,k denotes the observation data of the radar on the qth target at the kth continuous time, h(ξ q,k|k-1 ) denotes the predicted observation data obtained by the measurement function of the predicted state of the qth target at the kth continuous time;
[0079] The state estimation information of the qth target at the kth continuous time is obtained according to the innovation between the observation data and the prediction data, the Kalman gain matrix and the prediction state information of the qth target at the kth continuous time.
[0080] Compared with the prior art, the present application has the beneficial effects that:
[0081] The radar multi-dimensional resource management method provided by the present application firstly obtains the prediction state information of the qth target at the kth continuous time based on the state estimation information of the qth target at the k-1th continuous time and the target state transition matrix, then introduces the jammer information into the active anti-jamming MDP model, and combines the prediction state information of the qth target at the kth continuous time to obtain the radar MRM scheme model to be solved at the kth continuous time, after that, the radar MRM scheme model to be solved at the kth continuous time is decomposed into a first sub-problem related to the transmission frequency and a second sub-problem related to the transmission power, and the MRM scheme is obtained according to the optimal transmission scheme and the random transmission scheme determined by the first sub-problem and the second sub-problem, so as to obtain the observation data based on the radar executing the MRM scheme, and determine the state estimation information of the qth target at the kth continuous time, and then enter the k+1th continuous time until the requirement of stopping operation is reached. The present application integrates the radar resource management and the active anti-jamming into a coherent framework, can reduce the dynamic jamming and improve the tracking performance at the same time, has high overall design comprehensiveness, and is more consistent with the mode requirement of actual application.
[0082] The present application will be further described in detail below with reference to the accompanying drawings and the present application. BRIEF DESCRIPTION OF DRAWINGS
[0083] Figure 1 is a flowchart of a radar multi-dimensional resource management method for multi-target tracking in a dynamic electromagnetic environment provided by an embodiment of the present application;
[0084] Figure 2 is a flowchart of a radar multi-dimensional resource management method for multi-target tracking in a dynamic electromagnetic environment provided by an embodiment of the present application;
[0085] Figure 3 is a deployment diagram of a target relative to a radar provided by an embodiment of the present application;
[0086] Figure 4 is a comparison diagram of MTT performance parameter curves of each scheme in an interference environment 1 provided by an embodiment of the present application;
[0087] Figure 5 is a comparison diagram of MTT performance parameter curves of each scheme in an interference environment 2 provided by an embodiment of the present application. DETAILED DESCRIPTION
[0088] The application will be described in further detail below with reference to the embodiments. However, the embodiments of the application are not limited thereto.
[0089] Embodiment one
[0090] Please refer to Figure 1 and Figure 2 , Figure 1 is a flowchart of a radar multi-dimensional resource management method for multi-target tracking in a dynamic electromagnetic environment provided by an embodiment of the application, Figure 2 is a flowchart of a radar multi-dimensional resource management method for multi-target tracking in a dynamic electromagnetic environment provided by an embodiment of the application. The application provides a radar multi-dimensional resource management method for multi-target tracking in a dynamic electromagnetic environment. The radar multi-dimensional resource management method comprises the following steps:
[0091] Step 1: obtaining jammer information of a jammer, the jammer information comprising a frequency and a jamming mode of the jammer.
[0092] In this embodiment, initialization is first performed. The initialization is as follows: k represents the kth continuous time, k is a positive integer, and the initial value of k is 1; q represents the target (such as a UAV) number, 1≤q≤Q and q is a positive integer, Q is the total number of targets, and a jammer is arranged on each target, thereby obtaining the frequency and jamming mode of the jammer.
[0093] Specifically, there is an FA radar located at coordinates (x, y) and Q independent point targets moving in a two-dimensional space. All targets have been confirmed and tracked by the radar, so the number of these targets is known a priori. Each target is equipped with a jammer that can emit frequency-seeking and blocking jamming signals to reduce the tracking performance of the radar. The FA radar can change the carrier frequency of its transmitted signal to avoid the presence of jamming signals in the received echo. The qth target is initially located at (x q,0 ,y q,0 ), with an initial velocity where T is the time interval between consecutive time instants. The initial time is 0, and at time kT, the position and velocity of the qth target are represented as (x q,k ,y q,k ) and In this embodiment, the position coordinates of the radar are set as (0, 0) km, the number of targets Q is 3, the time interval T is 0.2 s, and the position and velocity parameters of the targets at the initial time are shown in Table 1.
[0094] Table 1
[0095] Target number 1 2 3 Position (km) (6,5) (8,6) (-3,3) Speed (m / s) (-4,2) (-4,-4) (2,6)
[0096] For the qthtarget, the transmitted signal of the FA radar at the kthconsecutive time instant can be represented as:
[0097]
[0098] wherein, represents the transmitted signal of the radar for the qthtarget at the kthconsecutive time instant, t is time, represents the transmitted power of the radar for the qthtarget at the kthtime instant, e q,k (t) represents the normalized complex envelope of the transmitted signal, f n represents the nthcarrier frequency, 1£n£N, N represents the number of available frequencies of the radar, f n =f1+(n-1)Df, f1represents the minimum carrier frequency that can be selected by the radar, Dfrepresents the step size between two consecutive carrier frequencies, Df>0.
[0099] Based on the above, the received signal after frequency down-conversion can be represented as:
[0100]
[0101] wherein, represents the received signal after frequency down-conversion at the kthconsecutive time instant, a q,k represents the attenuation of the signal strength, D q,k is the distance between the qthtarget and the radar, t q,k represents the target time delay, v q,k represents the Doppler shift relative to the carrier frequency f n , w q,k (t) represents the zero-mean complex Gaussian noise.
[0102] For example, it is assumed that the number of available transmitted frequencies N of the radar is 3, the initial carrier frequency f1is 1.4 GHz, and the step size Df between two consecutive carrier frequencies is 100 MHz.
[0103] Step 2, based on the state estimation information of the qthtarget at the k-1thconsecutive time instant and the target state transition matrix, the predicted state information of the qthtarget at the kthconsecutive time instant is obtained.
[0104] Specifically, the motion of all the tracked targets can be described by the nearly constant velocity (NCV) model and the nearly coordinated turn (NCT) model with two different reverse rotation factors. For the qthtarget, its dynamic motion can be represented as:
[0105] ξ q,k =Fξ q,k-1+u q,k-1 (3)
[0106] wherein, ξ q,k denotes the state of the qth target at the kth consecutive time, [·] T denotes the transpose of a matrix or a vector, F is a target state transition matrix, ξ q,k-1 denotes the state of the qth target at the k-1th consecutive time, u q,k-1 denotes the process noise of the qth target at the k-1th consecutive time.
[0107] For the NCV model, F is denoted as F NCV , and the target state transition matrix is denoted as:
[0108]
[0109] wherein, denotes the Kronecker operator, and I2 denotes a 2-order unit matrix.
[0110] For the NCT model, F is denoted as F NCT , and the target state transition matrix is denoted as:
[0111]
[0112] wherein, w denotes the rotation angular velocity.
[0113] The process noise u q,k-1 is zero-mean Gaussian noise with a known covariance Q q,k-1 , and the expression of Q q,k-1 is:
[0114]
[0115] wherein, q0 denotes the process noise intensity, q030, which controls the uncertainty level of the motion state of the target. In this embodiment, q0 is set to 0.01, for example, the motion models of the targets 1 and 2 are NCV throughout the simulation of the 2000th consecutive time, and the motion model of the target 3 includes NCV and NCT, and will rotate at an angular velocity of w=-8° / s between the 1000th-1200th consecutive time.
[0116] In this embodiment, the predicted state information of the qth target at the kth consecutive time is denoted as: ξ q,k|k-1 =Fξ q,k-1|k-1 , wherein, ξ q,k|k-1 denotes the predicted state information of the qth target at the kth consecutive time, and ξ q,k-1|k-1 denotes the state estimation information of the qth target at the k-1th consecutive time.
[0117] In addition, the following is introduced for the convenience of understanding the subsequent steps.
[0118] To degrade the tracking performance of a radar, a jammer mounted on a target attempts to mask the target information in the radar return by transmitting a jamming signal. Whether a FA radar is affected by a jammer depends mainly on the overlap of the available carrier frequencies of the radar and the jammer. When the available frequency range of the radar is completely covered by the jammer, the jamming is most effective, otherwise the radar can avoid the jamming by transmitting signals on the uncovered frequencies. The present invention concerns the case where the jammer operates in a frequency range larger than the radar frequency range. The jammer can use two types of jamming modes: a spot jamming mode and a barrage jamming mode. In the spot jamming mode, the jammer transmits a narrowband signal, while in the barrage jamming mode, the jammer transmits a wideband signal. Due to the limited peak power of transmission, the power density of the barrage jamming is usually lower than that of the spot jamming, which can result in a degraded jamming performance. Let P J,q denote the jammer signal power on the qth target, then the power density of the jamming signal can be expressed as:
[0119]
[0120] where B J,r denotes the jamming signal bandwidth. For the spot jamming, B J,r is slightly larger than the radar signal bandwidth; for the barrage jamming, B J,r depends on the number of carrier frequencies covered by the jammer. For example, if the jammer wants to jam the radar on the carrier frequency f i , it can use a spot jamming with B J,r ≈ Df; if the jammer wants to jam f i and f j at the same time, it needs to use a barrage jamming with bandwidth B J,r = MDf, where M = |i - j| + 1 denotes the number of frequencies between f i and f j . In this embodiment, it is specified that for any qth target, the jammer signal power P J,q on it is 40 W, when using the spot jamming mode, the jamming signal bandwidth B J,r is 2 MHz, while when using the barrage jamming mode, the jamming signal bandwidth B J,r is set according to the specific jamming strategy.
[0121] At the kth consecutive time instant, the measurement of the qth target has the following form:
[0122] z q,k = h(ξ q,k ) + w q,k (8)
[0123] where h(·) is a measurement function, which is specifically expressed as:
[0124]
[0125] where D q,k , q q,k and v q,k correspond to different measurement components, i.e., range, azimuth and Doppler, respectively, and their specific expressions are as follows:
[0126]
[0127] where w q,k represents an observation noise, which is a zero-mean Gaussian noise with a covariance matrix S q,k , and the expression of S q,k is:
[0128]
[0129] where and are the Cramer-Rao Lower Bound (CRLB) of the estimation of the Mean-Square Error (MSE) of the target range, azimuth and velocity, respectively, and they satisfy the following relationship:
[0130]
[0131] where T R,d represents the corresponding dwell time, B R,NN represents the 3dB receiving beam width. L q,k represents the SINR of the qth target at the kth continuous time, and its expression is:
[0132]
[0133] where g n,q represents the Radar Cross Section (RCS) of the qth target at the carrier frequency f n , G R,t and G R,r represent the gains of the transmitting antenna and the receiving antenna, respectively, l R,t represents the signal wavelength, N R is the number of pulses in a Coherent Processing Interval (CPI), G R,p represents the processing gain, k0 represents the Boltzmann constant, T N represents the noise temperature of the radar, and B R,rdenotes the bandwidth of the radar received signal, F R,r denotes the noise figure, denotes the transmit power of the radar for the qth target at the kth continuous time. is the interference power into the radar, which is expressed as:
[0134]
[0135] where G J,t denotes the transmit antenna gain of the jammer, l J,t denotes the wavelength of the jammer, denotes the interference power density related to the bandwidth. The radar can force the jammer to increase its bandwidth B J,r by taking some anti-jamming measures, so that the interference power into the radar is reduced, and thus the SINR is increased, thereby improving the radar tracking performance. In addition, the RCS of the target varies at different frequencies. When the corresponding wavelength is close to the size of the target, the target will be in the Mie region, in which case the target will have a larger RCS for lower frequencies in some directions. Therefore, it can be considered that the RCS of the target decreases with the increase of the carrier frequency. In the present example, N R is set to 128, G R,t is 30 dB, G R,r is 30 dB, T R,d is 1 ms, G R,p is 30 dB, B R,r is 1 MHz, F R,r is 3 dB, B R,NN is 0.5°, G J,t,q is 3 dB for any q; the RCS of all targets is in the Mie region and remains unchanged over time, and changes with the transmit frequency, and for the three different transmit frequencies used by the radar, the RCS of each target is set to [3, 2, 1] m 2 , wherein the higher carrier frequency corresponds to the smaller RCS value.
[0136] Step 3, import the jammer information into the active anti-jamming MDP (Markov Decision Process) model, and combine the predicted state information of the qth target at the kth continuous time to obtain the radar MRM scheme model to be solved at the kth continuous time.
[0137] Specifically, in order to solve the multi-target detection problem in an unknown dynamic jamming environment, the present application regards the anti-jamming strategy design as a sequential decision problem, wherein the transmit frequency and the transmit power are decision variables. The MDP model is used to describe the decision problem, and the RL method is used to solve the problem. An MDP is defined by a tuple .
[0138] State of a finite set;
[0139] Action of a finite set;
[0140] State transition probability from current state s to next state s after performing action a
[0141] r: reward function that gives a scalar reward r(s, a, s') when the agent transitions to state s' from state s performing action a.
[0142] The general interaction between the agent and the environment is described as follows: at each time t, the agent receives information from the environment in the form of a state s t and selects an action a based on a policy t which is a function mapping states to probability simplices over the action space. After performing action a t , the state transitions to a new state s t+1 , and the agent receives a reward r t = r(s t , a t , s t+1 ). In the framework designed in this invention, the radar is the agent and all the jammers together constitute the jamming environment. At each time t, assume that the radar is in state s t and takes action a t ; then, the radar receives the jamming signal, accordingly, the state transitions to s t+1 , and the radar obtains r t .
[0143] On the basis of the above, step 3 can specifically include:
[0144] Step 3.1, at the kth continuous time, the action taken by the radar on the qth target is obtained.
[0145] Specifically, the action of the radar on the qth target can be described as:
[0146]
[0147] Wherein, represents the action taken by the radar on the qth target at the kth continuous time, represents the carrier frequency of the signal transmitted by the radar on the qth target when taking action at the kth continuous time, denotes the action space of the radar, denotes the number of available carrier frequencies that the radar can select.
[0148] For the qthtarget at the kthconsecutive time, the action taken by the radar is The set of actions a of the radar for all targets at the kthconsecutive time k denotes the action space of the radar,
[0149]
[0150] In this embodiment,
[0151] Step 3.2, at the kthconsecutive time, obtain the action space of the jammer on the qthtarget.
[0152] Specifically, in ECCM, the action of the jammer can be perceived by the radar, and the result is regarded as the observation of the radar to form the state space of the radar. The action of the jammer is described as:
[0153]
[0154] wherein, denotes the action taken by the jammer on the qthtarget at the kthconsecutive time, denotes the number of available jamming actions, is the action space of the jammer, at the kthconsecutive time, the jammer on the qthtarget perceived by the radar is If , it indicates that the radar is not jammed, if , it indicates that the jammer releases a frequency-agile jamming signal with the frequency of , if , it indicates that the jammer releases a barrage jamming signal, is the jamming signal frequency emitted by the jammer on the qthtarget at the kthconsecutive time when taking the action .
[0155] Step 3.3, according to the action taken by the radar on the qthtarget and the action space of the jammer on the qthtarget, obtain the interaction information of the radar and the jammer on the qthtarget at the kthconsecutive time.
[0156] Specifically, for the intelligent jammer, the jamming strategy will use some historical actions of the radar to determine the action. Therefore, in order to successfully counter the jamming strategy, the radar also needs to make decisions according to the interaction information, which can be described as an alternating sequence, that is, the interaction information is represented as:
[0157]
[0158] in, This represents the interaction information between the radar and the jammer on the q-th target at the k-th consecutive time step. That is, a set of observations.
[0159] Step 3.4: At the k-th consecutive time, based on the last M sets of observations in the interaction information between the radar and the jammer on the q-th target at the k-th consecutive time, obtain the radar's state against the q-th target at the k-th consecutive time.
[0160] Specifically, due to The length of the interval increases over time and cannot be directly used as the state; therefore, the last M observations are selected to approximate the interval. Therefore, the state space of the radar can be represented as:
[0161]
[0162] in, For the state space of the radar, For the first in the radar state space One possible state, The number of states,
[0163] Based on the given state space Introducing symbols Let represent the state of the radar at the k-th consecutive time step towards the q-th target, written as This represents the action taken by the q-th target at the kM consecutive time points. This represents the action taken by the jammer on the q-th target at the kM-th consecutive time. This represents the action taken by the q-th target at the (k-1)th consecutive time step. This represents the action taken by the jammer on the q-th target at the (k-1)-th consecutive time. For all targets, the radar state at time k is represented as:
[0164]
[0165] In this embodiment, based on the guidelines established above, the following settings are made: The value is 4.
[0166] Step 3.5: When the radar takes action against the q-th target Then, based on the maximum likelihood method, the final reward function is obtained according to the sum of all rewards obtained by the radar from the state of the qth target at the kth continuous moment to the state of the qth target at the (k+1)th continuous moment and the number of transitions from the state of the qth target at the kth continuous moment to the state of the qth target at the (k+1)th continuous moment.
[0167] Specifically, in RL, the reward is defined according to the goal of the agent. Similarly, for the radar system, a suitable criterion is needed to guide the radar against interference. SINR can effectively capture the impact of interference on the radar, so in the present application, it is taken as the main component of the reward to help the radar find the best anti-interference strategy. According to equation (13), the SINR of the radar for the qth target at the kth continuous moment can be calculated as:
[0168]
[0169] wherein, RCSq(k) represents the radar cross section (RCS) of the qth target at the carrier frequency of the radar, Pj(k) represents the interference power received by the radar, which can be calculated by equation (14) using to emphasize that this power is a function of .
[0170] Generally speaking, the interference power entering the radar is usually much larger than the noise power, so the SINR can be simplified as:
[0171]
[0172] wherein, if then For the jammer, the interference signal bandwidth B R,r is related to the actions of the radar and the jammer. Therefore, according to equation (7), the interference power density is also a function of .
[0173] Since the target is moving and the transmit power is adjustable, the changes of both make the SINR dynamic for the same . When the target and the jammer reach new positions, the radar cannot use the past perceived SINR to predict the tracking performance, which brings difficulties to resource management. In order to solve this problem, the expression of SINR is processed in the present application to form a suitable reward function:
[0174]
[0175] wherein, D q,kkrepresents the distance estimation value between the qth target and the radar at the kth continuous time. In addition, the SINR estimated according to the received signal can fluctuate in practice. In order to eliminate this instability, the present application utilizes the maximum likelihood method to estimate the final reward function:
[0176]
[0177] wherein, represents the final reward function, represents the state of the radar for the qth target at the k+1th continuous time, represents the number of times of state transition of the radar from the state of the radar for the qth target at the kth continuous time to the state of the radar for the qth target at the k+1th continuous time, represents the sum of all rewards obtained by the state transition of the radar from the state of the radar for the qth target at the kth continuous time to the state of the radar for the qth target at the k+1th continuous time, and the expression thereof is:
[0178]
[0179] wherein, is the reward function at the bth calculation
[0180] if then
[0181] Step 3.6, based on the maximum likelihood method, the state transition probability is obtained according to the number of times of state transition of the radar from the state of the radar for the qth target at the kth continuous time to the state of the radar for the qth target at the k+1th continuous time.
[0182] Specifically, generally speaking, the strategy of the jammer is usually unknown to the radar, and the radar has no prior knowledge about the environment at the beginning. Therefore, the radar needs to learn the underlying dynamic characteristics first through interaction with the environment. The present application utilizes the maximum likelihood method to estimate the state transition probability, which is represented as:
[0183]
[0184] wherein, represents the state transition probability, the expression of is:
[0185]
[0186] and satisfies: if then
[0187] Step 3.7, predicting the Bayesian information matrix at the kth consecutive time according to the predicted state information of the qth target at the kth consecutive time, the final reward function and the state transition probability, so as to determine the radar MRM scheme model to be solved at the kth consecutive time according to the predicted Bayesian information matrix at the kth consecutive time.
[0188] In one specific embodiment, step 3.7 can include:
[0189] Step 3.71, predicting the Bayesian information matrix at the kth consecutive time according to the predicted state information of the qth target at the kth consecutive time, the final reward function and the state transition probability.
[0190] Specifically, in multi-target tracking, the BCRLB (Bayesian Cramer-Rao Lower Bound) is used to constrain the error variance of target state estimation, and is predicted in the tracking recursive loop. Therefore, the predicted BCRLB is usually used to quantify the target tracking performance. The present application combines the BCRLB and introduces an anti-interference performance metric, thereby forming an MRM scheme framework that can jointly adjust the transmission frequency and the transmission power. As the basis of the MRM scheme, the analytical expression of the predicted Bayesian information matrix (BIM) at the kth consecutive time is:
[0191]
[0192] wherein, denotes the Bayesian information matrix at the kth consecutive time, denotes the Bayesian information matrix at the (k-1)th consecutive time, denotes the transmission power of the radar for the qth target at the (k-1)th time, denotes the action taken by the radar for the qth target at the (k-1)th consecutive time, H q,k denotes the Jacobian matrix of the radar for the qth target at the kth consecutive time, S q.k denotes the covariance matrix of the observation noise of the radar for the qth target at the kth consecutive time, S q,k related to the SINR, denotes the final reward function, D q,kk-1 denotes the predicted value of the distance between the qth target and the radar at the kth consecutive time.
[0193] The present application also expects that the BIM can predict the trend of the jammer. According to the MDP principle set forth above, it can be seen that the state transition probability can predict the future state according to the current state and action, and the reward function can reflect the comprehensive performance of the radar in the future state to a certain extent. Therefore, the present application combines the above two parameters, and designs the expected SINR parameter as the SINR prediction parameter when the MRM scheme is optimized and solved, and the expression of the parameter is:
[0194]
[0195] The expression of the expected return is as follows:
[0196]
[0197] Only related to the action of the radar, and since the current state of the radar is known, the expression is omitted in the formula (29) and the formula (30).
[0198] Step 3.72, obtaining the predicted BCRLB according to the inverse of the Bayesian information matrix of the kth continuous time.
[0199] Specifically, the in the formula (29) will be used to calculate the predicted BIM, and the predicted BCRLB is defined as the inverse of the predicted BIM, that is, the predicted BCRLB is expressed as:
[0200]
[0201] Wherein, represents the predicted BCRLB.
[0202] Step 3.73, obtaining the MTT index according to the predicted BCRLB.
[0203] Specifically, The diagonal elements of the matrix can reflect the lower limit of the target state estimation MSE, and the reward function also participates in the tracking performance evaluation in the form of the expected SINR. In order to quantitatively describe the target tracking accuracy, the following MTT index is defined:
[0204]
[0205] Wherein, represents the MTT index, and Tr represents the trace of the matrix.
[0206] Step 3.74, obtaining the target function of the radar MRM scheme model according to the MTT index.
[0207] Specifically, for the MTT performance, the worst target tracking accuracy is used as the objective function of the MRM scheme optimization model, and the objective function of the radar MRM scheme model is represented as:
[0208]
[0209] wherein, represents the objective function of the radar MRM scheme model, P R,k represents a set of the transmission power of the radar for all targets at the kth continuous moment, a k represents a set of actions of the radar for all targets at the kth continuous moment, a k The meaning of a
[0210] Step 3.75, determining the radar MRM scheme model to be solved at the kth continuous moment according to the objective function of the radar MRM scheme model, and the radar MRM scheme model to be solved at the kth continuous moment is represented as:
[0211]
[0212] wherein, the constraint set The expression of
[0213]
[0214] wherein, P Total represents the total amount of power resources that can be allocated by the radar, P min represents the lower limit of the power resources that can be allocated by the radar for a single target, P max represents the upper limit of the power resources that can be allocated by the radar for a single target. The constraint set indicates that the total resources of the radar used for MTT are limited. In the embodiment, P Total is set to 5kW, P min is set to 0.05P Total , P max is set to P Total .
[0215] Step 4, decomposing the radar MRM scheme model to be solved at the kth continuous moment into a first sub-problem related to the transmission frequency and a second sub-problem related to the transmission power, and obtaining the MRM scheme according to the optimal scheme and the random scheme determined according to the first sub-problem and the second sub-problem.
[0216] In a specific embodiment, step 4 can include:
[0217] Step 4.1, decomposing the radar MRM scheme model to be solved at the kth continuous moment into a first sub-problem related to the transmission frequency and a second sub-problem related to the transmission power.
[0218] Specifically, the optimization problem corresponding to equation (34) is decomposed into a first subproblem concerning the transmission frequency and a second subproblem concerning the transmission power. The first subproblem and the second subproblem are expressed as follows:
[0219]
[0220]
[0221] in, This represents the first subproblem. This represents the second subproblem, (·). * This represents the optimal solution obtained for the corresponding parameters.
[0222] The solution algorithm designed in this invention solves the two subproblems sequentially, and the obtained optimal solution... It has been proven to be the optimal solution to the original optimization problem.
[0223] Step 4.2: Using the RL-based policy iteration method, the new policy is updated according to the value function of the policy constructed by the Bellman equation, and the optimal solution of the first subproblem is obtained.
[0224] Specifically, for the subproblem corresponding to equation (36), this invention employs a policy iteration method based on RL for solution. The purpose of this solution method is to continuously update the state-to-action mapping policy in the MDP, and finally iterate to an optimal policy to adjust the radar state. Mapping to the optimal solution In the design of this invention, strategy p q,k value function Given by the Bellman equation:
[0225]
[0226] in, The expression for U(·) is referenced in equation (30). Indicates strategy p q,k The value function, Indicates a radar status Mapping to Action The strategy is defined by g∈[0,1], where g is the discount factor, representing the preference of the current reward for future rewards. If γ is 0, it means the agent only cares about immediate returns; if γ is close to 1, it means the agent is more concerned with long-term future returns. In practical applications, future returns are very important for the overall mission of the radar, so introducing the discount factor γ can improve the stability of the radar's operation throughout the mission.
[0227] Based on equation (38), the expression for the new policy obtained through policy iteration can be given:
[0228]
[0229] New strategy p' q,k A new value function can be generated For the iteration to continue to obtain a new strategy. The above process is repeated until the strategy can no longer be improved, at which time the strategy can be considered to be the optimal strategy, that is Thus the optimal solution of the first sub-problem can be obtained:
[0230]
[0231] Combine all According to the fixed format, that is
[0232] Step 4.3, based on the gradient projection method, the optimal solution of the second sub-problem is obtained.
[0233] Specifically, for the sub-problem corresponding to formula (37), since the problem is convex, an effective convex optimization algorithm can be flexibly selected to obtain the optimal solution. In this embodiment, the gradient projection method is selected as the solving algorithm for the sub-problem.
[0234] Step 4.4, the MRM scheme is obtained according to the optimal emission scheme and the random emission scheme determined by the first sub-problem and the second sub-problem.
[0235] Specifically, in the actual task process, the interaction between the radar and the target is usually carried out online. For an unknown interference environment, the radar must select a new action to obtain the information of the state first perceived in the online interaction process. For the states that have been visited, the radar is still in a dilemma: is it to use the obtained knowledge to form an MRM scheme to generate the best emission scheme, or is it to explore the environment to discover a better emission action of the MRM. This is a classic exploration and exploitation problem. Therefore, the radar needs to find a balance point between detection and exploitation, especially in the online tracking process under an unknown interference environment. It is assumed that at time k, the radar is in state s k This is divided into two possible cases:
[0236] (1) s k is a state that has not been visited: in this case, there is at least one target state is first observed by the radar. In this new state, the radar does not know the transition probability, that is Therefore, it is necessary to detect the outside. Accordingly, in the present application, it is stipulated that the radar selects a random emission scheme to detect the interference environment, wherein is a set of emission frequencies selected from a uniform distribution, is a set of transmit powers selected from a uniform distribution of total resources.
[0237] (2) s k is a state that has been visited: in this case, each target state has been observed by the radar at least once in the past, i.e. Since the transition probabilities of s k are estimated from past interactions, inaccuracies can occur. Therefore, the radar needs to decide whether to exploit the MRM scheme to generate the best action to improve the MTT performance or to discover new actions to provide new knowledge for the MRM. The present invention uses an e-greedy algorithm to solve this dilemma, which stipulates that the radar has a probability of e (0£e£1) to use a random scheme to explore the environment, while has a probability of 1-e to use the optimal scheme generated by the MRM scheme represents the optimal transmit scheme, represents a set of optimal solutions of the first sub-problem for all targets, represents the optimal solution of the second sub-problem.
[0238] Therefore, the MRM scheme is represented as:
[0239]
[0240] where e is the greed rate, 0£e£1.
[0241] Step 5, based on the observation data obtained by the radar executing the MRM scheme, determine the state estimation information of the qth target at the kth continuous time.
[0242] Specifically, the radar executes the MRM scheme obtained in step 4, i.e. the radar observes each target and imports the observation data into the tracking processor for processing to obtain the estimation information of each target at the current kth continuous time. Specifically, the tracking processor of the radar utilizes the observation data of the target obtained by observation, and then utilizes the Kalman filtering method for filtering processing to obtain the state estimation information of the target at the current time, and then performs state prediction of the target at the next time, etc. The present invention does not have specific requirements for the above filtering method, which can be flexibly selected according to the environmental characteristics and task requirements.
[0243] In one specific embodiment, step 5 can include:
[0244] Step 5.1, the radar executes the MRM scheme to obtain the observation data z q,k of the qth target at the kth continuous time by the radar, the specific expression can be referred to formula (8).
[0245] Step 5.2, obtain the innovation between the observation data and the predicted data according to the measurement function of the observation data of the qth target at the kth continuous time instant by the radar and the predicted state information of the qth target at the kth continuous time instant, and the innovation between the observation data and the predicted data is represented as:
[0246] v q,k q,k -h(ξ q,k|k-1 ) (42)
[0247] wherein, z q,k represents the observation data of the qth target at the kth continuous time instant by the radar, h(ξ q,k|k-1 ) represents the predicted observation data of the qth target at the kth continuous time instant by the predicted state through the measurement function, and the specific expression can be referred to formula (9), which will not be repeated here.
[0248] Step 5.3, obtain the state estimation information of the qth target at the kth continuous time instant according to the innovation between the observation data and the predicted data, the Kalman gain matrix and the predicted state information of the qth target at the kth continuous time instant.
[0249] Specifically, the Kalman gain matrix is represented as:
[0250]
[0251] P q,k|k-1 q,k-1|k-1 F T (44)
[0252]
[0253] wherein, P q,k-1|k-1 represents the covariance matrix of the conditional error of the radar for the qth target at the kth continuous time instant, H q,k represents the Jacobian matrix of the radar for the qth target at the kth continuous time instant.
[0254] Therefore, the state estimation information of the qth target at the kth continuous time instant is represented as:
[0255] ξ q,k|k q,k|k-1 +K q,k v q,k (46)
[0256] wherein, ξ q,k|k represents the state estimation information of the qth target at the kth continuous time instant.
[0257] Step 6, return to step 1, enter the k+1th continuous time instant, and continue until the requirement of stopping operation is reached.
[0258] Specifically, let k=k+1, and return to step 1, that is, the radar enters the flow of the next moment, until the overall task meets the requirement of stopping running. The requirement of stopping running can be that the tracking accuracy meets a certain requirement, or the time of executing the task reaches a specified value, etc. In the present example, it is specified that the radar executes the task until the time reaches 2000.
[0259] The present application builds an active anti-jamming MDP model based on the RL framework to counter dynamic jamming. In order to describe the anti-jamming problem, the present application makes the state space of the MDP contain the interaction information between the radar and the jammer, and the action space is the transmission frequency of the radar; by designing a special reward signal, the influence of the jammer on the radar is quantified; the state transition function is used to represent the dynamic characteristics of the environment, which is learned by interacting with the unknown jamming environment. Based on the active anti-jamming model, the present application introduces its return and state transition function to calculate the SINR, and further forms a predicted BCRLB as a measure to evaluate the target tracking accuracy under the jamming environment. In this case, the radar can change the value of the performance index by adjusting the transmission frequency and power, and then determine the target tracking performance. For a given performance measure, the MRM scheme of the present application is considered as an optimization problem, in which the frequency and power parameters are coupled into the objective function and constraints. This optimization problem is proved to be decoupled by analysis, so as to be equivalently recast as two sub-problems, one for countering dynamic jamming and the other for power allocation. In order to solve the decoupling problem, the present application develops a two-step solution method to obtain a joint transmission scheme: in the first step, a RL method is used to select the optimal transmission frequency to counter dynamic jamming; in the second step, the power allocation problem is solved by an optimization algorithm. The obtained joint transmission scheme is strictly proved to be a globally optimal solution. In addition, the whole MTT process is considered as an online learning process, and the radar needs to use the MRM framework to improve the MTT performance while acquiring new knowledge to update the state transition function. Therefore, the present application uses an e-greedy algorithm to balance the relationship between model updating and MRM scheme utilization until the online tracking process ends.
[0260] In order to counter dynamic jamming, the present application designs an active anti-jamming Markov decision process model based on RL; by calculating the SINR and Bayesian Cramer-Rao lower bound under the jamming environment, a performance optimization model of multi-target tracking is established; the MRM problem is decoupled into two sub-problems, one for countering dynamic jamming and the other for resource allocation, and based on decoupling analysis, a two-step solution technique is proposed to solve the optimization sub-problems resulting therefrom, and the obtained solution is proved to be able to maximize the target tracking performance. The simulation results show that under a given power budget, this method can significantly suppress dynamic jamming and improve target tracking accuracy.
[0261] The effect of the present application is further verified and illustrated by the following simulation.
[0262] 1. Simulation settings:
[0263] This simulation covers 2000 time points of radar MTT (Multi-Track To-Track) missions. The radar is located at coordinates (0,0) km, the number of available transmission frequencies N is set to 3, the initial carrier frequency f1 is set to 1.4 GHz, the step size Df between two subcarriers is set to 100 MHz, the tracking interval T is set to 0.2 s, the process noise intensity q0 is set to 0.01, and the total power resource P... Total Set to 5kW, P min Set to 0.05P Total P max Let P be the value of P. Total More radar-related parameters are shown in Table 2.
[0264] Table 2
[0265] Parameter <![CDATA[N R ]]> G R,t ]]> G R,r ]]> [CAT R,d ]]> G R,p ]]> B R,r ]]> F R,r ]]> B R,NN ]]> Set value 128 30 dB 30 dB 1 ms 30 dB 1 MHz 3 dB 0.5°
[0266] The number of targets, Q, is set to 3. The RCS of all targets is located within the Mie region and remains constant over time, changing only with the transmission frequency. For the three different transmission frequencies used by the radar, the RCS of each target is set to [3, 2, 1]m. 2 Higher carrier frequencies correspond to smaller RCS values. Table 3 provides the initial states of all targets. For the simulated 2000 time points, the motion models for targets 1 and 2 are NCT, while target 3 will rotate at an angular velocity of w = -8° / s between time points 1000 and 1200, therefore its motion model includes both NCT and NCV. The angular spans of these targets associated with the radar system are as follows: Figure 3 As shown.
[0267] Table 3
[0268] Target number 1 2 3 Position (km) (6,5) (8,6) (-3,3) Speed (m / s) (-4,2) (-4,-4) (2,6)
[0269] Each target is equipped with a self-protection jammer, so there are a total of 3 jammers. Assume that each jammer has the same parameters, as shown in Table 4.
[0270] Table 4
[0271]
[0272] Furthermore, the radar's historical actions will serve as crucial information for these jammers. In this simulation, the jammers can utilize the acquired historical actions to formulate the following three strategies:
[0273] Jamming Strategy 1: The jammer uses actions and previous actions To jointly determine the jamming signal to be transmitted. When When the radar is not jammed, the jammer transmits a frequency-agile jamming signal with carrier frequency ; otherwise, the jammer transmits a barrage jamming signal with bandwidth
[0274] Jamming strategy 2: The jammer determines the jamming signal to be transmitted currently by combining the current action and the previous actions . When the radar is not jammed, the jammer transmits a frequency-agile jamming signal with carrier frequency ; otherwise, the jammer transmits a barrage jamming signal with bandwidth
[0275] Jamming strategy 3: The jammer adopts a random jamming strategy and determines the jamming signal by combining historical information. When the radar is not jammed, the jammer transmits a frequency-agile jamming signal with carrier frequency ; otherwise, the jammer adopts a hybrid jamming mode, with a 20% probability of transmitting a frequency-agile jamming signal and an 80% probability of transmitting a barrage jamming signal.
[0276] Two jamming environments are set in the simulation to test the performance of the scheme:
[0277] Jamming environment 1: Jammer 1 and jammer 2 adopt jamming strategy 1, and jammer 3 adopts jamming strategy 2. The available carrier frequency of jammer 2 does not intersect with the carrier frequency of the radar, and the other jammer completely covers the available frequency of the radar.
[0278] Jamming environment 2: Jammer 1 adopts strategy 3, which completely covers the frequency of the radar. The other jammer adopts strategy 1, in which the available carrier frequency of jammer 3 only covers part of the radar frequency.
[0279] For the e-greedy algorithm, the greed rate is set to start from 1 when the radar encounters the above jamming environment, and decreases by 0.2 every 300 time points until it reaches 0.
[0280] Four benchmark schemes are set in the simulation to verify the superiority of the MRM scheme of the present application:
[0281] Random frequency and uniformly distributed resources (RFEDR): The radar evenly distributes the limited power resources to all targets and randomly selects a transmission frequency for tracking.
[0282] Anti-jamming and uniformly distributed resources (AEDR): The radar first collects interaction information (e = 1) from the jamming environment, collecting data for 300 time points (AEDR1) and 900 time points (AEDR2), respectively. After that, the total power resource is evenly distributed to all targets.
[0283] Random Frequency and Resource Management (RFRM): The radar utilizes existing resource management techniques and multi-target information from the tracking processor to obtain a power allocation scheme. Since the tracking processor can provide effective information throughout the tracking process, the radar can always perform these resource management schemes. Meanwhile, the radar selects a random frequency to illuminate the target.
[0284] Anti-Jamming and Resource Management (ARM): The radar can not only apply the RL method to solve the MDP to generate a transmission frequency, but also allocate limited power resources with the help of existing resource management methods. Similar to the AEDR, the ARM scheme also has two stages, ARM1 and ARM2.
[0285] 2. Simulation content:
[0286] The present application compares and simulates the MTT performance of the RFEDR scheme, the AEDR scheme, the RFRM scheme, the ARM scheme and the MRM scheme of the present application in two set interference environments.
[0287] 3. Analysis of simulation results:
[0288] It can be seen from Figure 4 and Figure 5 Compared with other schemes, whether it is BCRLB data or MSE data, the MRM scheme proposed in the present application always has lower corresponding data values in two interference environments, and can be considered to have better tracking performance and adaptability to interference environments.
[0289] In summary, the simulation experiment verifies the correctness, effectiveness and reliability of the present application.
[0290] The present application is aimed at radar anti-jamming, and on the basis of the RL framework, a radar active anti-jamming MDP model is constructed, the quantization of the radar active anti-jamming scheme is realized, and the effect of radar anti-jamming can be more intuitively reflected. The RL technology also enables the radar to normally perform anti-jamming work in an unknown interference environment, and to a certain extent, improves the anti-jamming effect.
[0291] The radar MRM scheme designed in the present application integrates resource management and active anti-jamming into a coherent framework, can simultaneously reduce dynamic interference and improve tracking performance, has high overall design comprehensiveness, and is more in line with the mode requirements of practical application. Compared with other schemes, the scheme designed in the present application has certain performance advantages.
[0292] The problems given by the MRM scheme are decoupled into two sub-problems, the solving difficulty is reduced, different targeted solving algorithms are used for different sub-problems, the adaptability of the algorithm to the problem is strengthened, and the overall solving efficiency and effect of the algorithm are further improved. The e-greedy algorithm is used to control the overall online learning process of the scheme, and the algorithm can better balance the comprehensive benefits of the radar exploration unknown interference environment strategy and the active anti-interference strategy, which further improves the overall execution effect and stability of the task.
[0293] Embodiment two
[0294] The embodiment provides an electronic device. The electronic device comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory are in communication with each other through the communication bus.
[0295] The memory is configured to store a computer program.
[0296] The processor is configured to execute the computer program to implement all or part of the steps of the radar multi-dimensional resource management method embodiment, and the specific implementation principles and technical effects are similar, and will not be repeated here.
[0297] Embodiment three
[0298] The embodiment provides a computer readable storage medium, which stores a computer program, and all or part of the steps of the radar multi-dimensional resource management method embodiment are implemented when the computer program is executed by a processor, and the specific implementation principles and technical effects are similar, and will not be repeated here.
[0299] It should be noted that the terms "first", "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features limited by "first", "second" can explicitly or implicitly include one or more features. In the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise specifically limited.
[0300] The above is a further detailed description of the present application in combination with specific preferred embodiments, and the specific implementation of the present application cannot be limited to these descriptions. For ordinary skilled persons in the art to which the present application belongs, without departing from the concept of the present application, a number of simple deductions or replacements can be made, and all should be regarded as falling within the protection scope of the present application.
Claims
1. A radar multi-dimensional resource management method for multi-target tracking in a dynamic electromagnetic environment, characterized in that, include: Step 1: Obtain jammer information, including the jammer's frequency and jamming mode; Step 2: Based on the state estimation information of the qth target at the (k-1)th consecutive time and the target state transition matrix, obtain the predicted state information of the qth target at the kth consecutive time, 1≤q≤Q, where Q represents the total number of targets; Step 3: Import the jammer information into the active anti-jamming MDP model, and combine it with the predicted state information of the qth target at the kth consecutive time to obtain the radar MRM scheme model to be solved at the kth consecutive time. Step 4: Decompose the radar MRM scheme model to be solved at the kth consecutive time into a first sub-problem related to the transmission frequency and a second sub-problem related to the transmission power, and obtain the MRM scheme based on the optimal transmission scheme and random transmission scheme determined by the first sub-problem and the second sub-problem. Step 5: Based on the observation data obtained by the radar executing the MRM scheme, determine the state estimation information of the q-th target at the k-th consecutive time step; Step 6: Return to step 1 and proceed to the (k+1)th consecutive time step until the requirement to stop running is met.
2. The radar multi-dimensional resource management method according to claim 1, characterized in that, The predicted state information of the q-th target at the k-th consecutive time step is represented as follows: x q,k|k-1 =Fξ q,k-1|k-1 Where, ξ q,k|k-1 ξ represents the predicted state information of the q-th target at the k-th consecutive time step. q,k-1|k-1 Let F represent the state estimation information of the q-th target at the (k-1)-th consecutive time step, and let F represent the target state transition matrix. For the NCV model, F is denoted as Fq. NCV For the NCT model, F is denoted as F0. NCT Target state transition matrix F NCV and the target state transition matrix F NCT They are represented as follows: Where I2 represents the second-order identity matrix, Let represent the Kronecker operator, T represent the time interval between consecutive moments, and ω represent the rotational angular velocity.
3. The radar multi-dimensional resource management method according to claim 1, characterized in that, Step 3 includes: At the k-th consecutive time interval, the action taken by the radar for the q-th target is acquired, and the action taken by the radar for the q-th target is expressed as: in, This represents the action taken by the radar against the q-th target at the k-th consecutive time. This indicates that the radar takes action at the k-th consecutive time. Let f1 be the carrier frequency of the signal transmitted to the q-th target, f1 be the minimum carrier frequency selectable by the radar, and Δf be the step size between two consecutive carriers, where Δf >
0. Indicates the radar's operational space. This indicates the number of available carrier frequencies that the radar can select; At the k-th consecutive time, the action space of the jammer on the q-th target is obtained, and the action of the jammer on the q-th target is represented as: in, This represents the action taken by the jammer on the q-th target at the k-th consecutive time. Indicates the number of available interference actions. Let be the action space of the jammer. At the k-th consecutive time, if This indicates that the radar has not been jammed. This indicates that the jammer released a frequency of Frequency-guided interference signals, if This indicates that the jammer has released a blocking jamming signal. For the jammer on the q-th target at the k-th consecutive time, when it takes action The frequency of the interference signal emitted at that time; Based on the radar's actions against the q-th target and the action space of the jammer on the q-th target, the interaction information between the radar and the jammer on the q-th target at the k-th consecutive time interval is obtained. This interaction information is expressed as: in, This represents the interaction information between the radar and the jammer on the q-th target at the k-th consecutive time step; At the k-th consecutive time interval, based on the last M sets of observations in the interaction information between the radar and the jammer on the q-th target at the k-th consecutive time interval, the state of the radar towards the q-th target at the k-th consecutive time interval is obtained, and the state of the radar towards the q-th target at the k-th consecutive time interval is expressed as: in, This represents the radar's state at the k-th consecutive time step for the q-th target. For the state space of the radar, This represents the action taken by the q-th target at the kM consecutive time points. This represents the action taken by the jammer on the q-th target at the kM-th consecutive time. This represents the action taken by the q-th target at the (k-1)th consecutive time step. This represents the action taken by the jammer on the q-th target at the (k-1)-th consecutive time. When the radar takes action against the q-th target Then, based on the maximum likelihood method, the final reward function is obtained by summing all the rewards obtained by the radar from the state of the radar at the kth consecutive time step for the qth target to the state of the radar at the (k+1)th consecutive time step for the qth target, and the number of transitions of the radar from the state of the radar at the kth consecutive time step for the qth target to the state of the radar at the (k+1)th consecutive time step for the qth target. Based on the maximum likelihood method, the state transition probability is obtained according to the number of transitions from the state of the radar at the k-th consecutive time step for the q-th target to the state at the (k+1)-th consecutive time step for the q-th target. The state transition probability is expressed as: in, Represents the state transition probability. And satisfy if but Based on the predicted state information of the q-th target at the k-th consecutive time step, the final reward function, and the state transition probability, the Bayesian information matrix at the k-th consecutive time step is predicted, so as to determine the radar MRM scheme model to be solved at the k-th consecutive time step based on the predicted Bayesian information matrix at the k-th consecutive time step.
4. The radar multi-dimensional resource management method according to claim 3, characterized in that, The final reward function is expressed as: in, This represents the final reward function. This represents the sum of all rewards obtained by the radar when transitioning from the state of the radar at the k-th consecutive time step to the state of the radar at the (k+1)-th consecutive time step for the radar targeting the q-th target. This represents the number of radar transitions from the state of the radar at the k-th consecutive time step targeting the q-th target to the state of the radar at the (k+1)-th consecutive time step targeting the q-th target. Let represent the state of the radar at the k-th consecutive time step for the q-th target, and if but The reward function for the b-th calculation reward function Represented as: in, This represents the estimated distance between the q-th target and the radar at the k-th consecutive time step. N represents the radar's transmit power for the q-th target at time k. R G represents the number of pulses in a coherent processing interval. R,t This indicates the gain of the transmitting antenna. Indicates the q-th target at the carrier frequency The radar cross section under the given conditions, λ R,t D represents the signal wavelength. q,k Let q be the distance between the target and the radar. G represents the power density of the interference signal. J,t λ represents the transmit antenna gain of the jammer. J,t B represents the wavelength of the jammer. R,r G represents the bandwidth of the radar received signal. R,r G represents the gain of the receiving antenna. R,p The processing gain is represented by k0, and the Boltzmann constant is represented by T. N F represents the noise temperature of the radar. R,r This represents the noise figure.
5. The radar multi-dimensional resource management method according to claim 3, characterized in that, Based on the predicted state information of the q-th target at the k-th consecutive time step, the final reward function, and the state transition probability, the Bayesian information matrix at the k-th consecutive time step is predicted. The radar MRM scheme model to be solved at the k-th consecutive time step is then determined based on the predicted Bayesian information matrix at the k-th consecutive time step, including: Based on the predicted state information of the q-th target at the k-th consecutive time step, the final reward function, and the state transition probability, predict the Bayesian information matrix at the k-th consecutive time step. The predicted BCRLB is obtained from the reciprocal of the Bayesian information matrix at the k-th consecutive time step. The MTT metric is obtained based on the predicted BCRLB, and the MTT metric is expressed as follows: in, This represents the MTT index, and Tr represents finding the trace of the matrix. Indicates the predicted BCRLB; The objective function of the radar MRM scheme model is obtained based on the MTT index, and the objective function of the radar MRM scheme model is expressed as follows: in, Let P represent the objective function of the radar MRM scheme model. R,k Let a represent the set of radar transmit power for all targets at the k-th consecutive time interval. k This represents the set of radar actions for all targets at the k-th consecutive time step; The radar MRM scheme model to be solved at the k-th consecutive time step is determined based on the objective function of the radar MRM scheme model. The radar MRM scheme model to be solved at the k-th consecutive time step is expressed as follows: in, P Total P represents the total amount of power resources that can be allocated to the radar. min P represents the lower limit of the power resources that a radar can allocate to a single target. max This indicates the upper limit of the power resources that the radar can allocate to a single target.
6. The radar multi-dimensional resource management method according to claim 5, characterized in that, The Bayesian information matrix at the k-th consecutive time step is represented as: in, Let Q represent the Bayesian information matrix at the k-th consecutive time step. q,k-1 This represents the covariance of the process noise at the k-th consecutive time step. I2 represents a 2-order identity matrix. Let represent the Kronecker operator, T represent the time interval between consecutive moments, q0 represent the process noise intensity, and F represent the target state transition matrix. This represents the Bayesian information matrix at the (k-1)th consecutive time step. Let represent the radar's transmit power for the q-th target at time k-1. This represents the action taken by the radar for the q-th target at the (k-1)-th consecutive time step. T H represents the transpose of a matrix or vector. q,k S represents the Jacobian matrix of the radar for the q-th target at the k-th consecutive time. q.k Let represent the covariance matrix of the radar's observation noise for the q-th target at the k-th consecutive time step. Indicates expected return. Let D represent the final reward function. q,k|k-1 B represents the predicted distance between the q-th target and the radar at the k-th consecutive time step. R,r B represents the bandwidth of the radar received signal. R,NN T represents the 3dB receive beamwidth. R,d Indicates the length of stay.
7. The radar multi-dimensional resource management method according to claim 5, characterized in that, Step 4 includes: The radar MRM scheme model to be solved at the k-th consecutive time step is decomposed into a first sub-problem concerning the transmission frequency and a second sub-problem concerning the transmission power. The first sub-problem and the second sub-problem are respectively expressed as: in, This represents the first subproblem. This represents the second subproblem, (·). * This represents the optimal solution obtained for the corresponding parameters; The policy iteration method based on RL obtains the optimal solution to the first subproblem by updating the new policy according to the value function of the policy constructed by the Bellman equation. Based on the gradient projection method, the optimal solution to the second subproblem is obtained; The MRM scheme is obtained based on the optimal launch scheme and random launch scheme determined by the first subproblem and the second subproblem. The MRM scheme is expressed as follows: in, Indicates a random launch scheme. This represents the set of transmission frequencies selected from a uniform distribution. This represents the set of transmission powers selected from a uniform distribution of total resources. This represents the optimal launch plan. Let represent the set of optimal solutions to the first subproblem of all objectives. Let e represent the optimal solution to the second subproblem, where e is the greed rate, and 0 ≤ e ≤ 1.
8. The radar multi-dimensional resource management method according to claim 1, characterized in that, Step 5 includes: The radar executes the MRM scheme to obtain the radar's observation data of the q-th target at the k-th consecutive time. The information between the observed data and the predicted data is obtained based on the measurement function of the radar's observation data of the q-th target at the k-th consecutive time and the predicted state information of the q-th target at the k-th consecutive time. This information is expressed as follows: v q,k =z q,k -h(ξ q,k|k-1 ) Among them, z q,k h(ξ) represents the radar's observation data of the q-th target at the k-th consecutive time. q,k|k-1 ) represents the predicted observation data obtained by the measurement function for the predicted state of the q-th target at the k-th consecutive time. Based on the information between the observed data and the predicted data, the Kalman gain matrix, and the predicted state information of the qth target at the kth consecutive time step, the state estimation information of the qth target at the kth consecutive time step is obtained.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing the computer program, implements the method according to any one of claims 1-8.
10. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1-8.
Citation Information
Patent Citations
Radar device
JP2010237129A
Tracking radar targets represented by multiple reflection points
US20210405178A1