A networking radar multi-target three-dimensional imaging resource scheduling method based on deep reinforcement learning

By constructing a resource scheduling model for multi-target 3D imaging of networked radar using deep reinforcement learning, the impact of radar node distribution on imaging quality is resolved, and efficient and accurate 3D target imaging is achieved under limited resources.

CN118981014BActive Publication Date: 2025-11-04AIR FORCE UNIV PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411005114.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2025-11-04
Estimated Expiration
2044-07-25

AI Technical Summary

Technical Problem

Existing networked radar multi-target three-dimensional imaging technology fails to effectively consider the impact of radar node distribution on imaging quality, resulting in unstable imaging quality and an inability to maximize the overall effectiveness of the radar system under limited resource conditions.

Method used

By employing a deep reinforcement learning-based approach, we construct a multi-target 3D imaging resource scheduling model for networked radars by perceiving target features and analyzing the impact of radar node distribution on imaging quality. We also design key elements of reinforcement learning to optimize radar resource allocation and ensure that imaging quality and resource consumption are minimized.

Benefits of technology

While ensuring imaging quality, the consumption of radar resources was optimized, the accuracy and efficiency of target 3D imaging were improved, and the probability of erroneous reconstruction was reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118981014B_ABST
    Figure CN118981014B_ABST
Patent Text Reader

Abstract

The application provides a networking radar multi-target three-dimensional imaging resource scheduling method based on deep reinforcement learning, and comprises the following steps: sensing target characteristics, calculating target required azimuth resolution; determining a two-dimensional imaging plane of each radar on the target and a projection relationship; analyzing the influence of radar node distribution on target three-dimensional imaging quality; constructing a networking radar multi-target three-dimensional imaging resource scheduling model; designing key elements of reinforcement learning; solving the networking radar multi-target three-dimensional imaging resource scheduling model, and performing three-dimensional imaging on the target according to the scheduling result. The best strategy of radar observation on each target can be obtained by solving the scheduling model, then each radar adopts a sparse aperture ISAR imaging algorithm to perform two-dimensional imaging on the target in the target set according to the scheduling strategy, finally, three-dimensional imaging is realized by combining a three-dimensional imaging method, and the three-dimensional imaging task of the radar on the multi-target is completed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of inverse synthetic aperture radar three-dimensional imaging, and particularly relates to a networking radar multi-target three-dimensional imaging resource scheduling method based on deep reinforcement learning. BACKGROUND

[0002] Traditional ISAR imaging technology can obtain high-quality two-dimensional images of targets, but it has problems of geometric distortion and difficulty in azimuth scaling, and the images obtained at different time periods and different angles differ greatly, which cannot extract accurate target features; ISAR three-dimensional imaging technology can further obtain height dimension information perpendicular to the imaging plane, and the target three-dimensional structure is unique, and the three-dimensional image is not sensitive to target attitude changes. Three-dimensional imaging of targets helps to accurately extract target features such as shape and size, and effective and stable features are very important for distinguishing different types of targets, so target three-dimensional imaging plays an important role in target classification, identification and other fields.

[0003] Current ISAR three-dimensional imaging technology mainly includes five different systems of three-dimensional imaging methods: single pulse and difference beam angle measurement ISAR three-dimensional imaging, ISAR three-dimensional imaging based on scattering center correlation, ISAR three-dimensional imaging based on imaging projection, InISAR three-dimensional imaging and array ISAR three-dimensional imaging. Among them, the ISAR three-dimensional imaging method based on imaging projection mainly selects multiple radars to observe the target at different angles, and then reconstructs the target three-dimensional structure based on the obtained target two-dimensional ISAR image and its projection relationship. This method has more scattering points and better imaging quality, and performs better than other algorithms.

[0004] However, this method requires multiple radars to image the target at different angles, and different radar distributions will have different effects on the quality of target three-dimensional imaging; and radar resources are often limited, and for multi-target three-dimensional imaging scenarios, we need to allocate appropriate radar resources for each target for imaging, and maximize the overall efficiency of the radar system under limited resources. However, the existing resource scheduling methods for networking radar multi-target three-dimensional imaging do not consider the influence of radar node distribution on imaging quality, so the scheduling results obtained by these methods cannot guarantee the imaging quality of the target. SUMMARY

[0005] In view of the deficiencies in the prior art, the present application provides a networking radar multi-target three-dimensional imaging resource scheduling method based on deep reinforcement learning, comprising the following steps:

[0006] Step 1: sensing the target features and calculating the required azimuth resolution of the target;

[0007] Second step: determine the two-dimensional imaging plane of each radar to the target and the projection relationship;

[0008] Third step: analyze the influence of radar node distribution on the quality of target three-dimensional imaging;

[0009] Fourth step: construct a multi-target three-dimensional imaging resource scheduling model for networked radars;

[0010] Fifth step: design the key elements of reinforcement learning;

[0011] Sixth step: solve the multi-target three-dimensional imaging resource scheduling model for networked radars, and perform three-dimensional imaging on the target according to the scheduling result.

[0012] Further, the first step is specifically:

[0013] Through the tracking results of each target in the detection area by the radar, the geographical position, motion speed, and distance and yaw angle between each target and each radar are obtained; then each radar transmits a small amount of pulses to perform two-dimensional ISAR coarse imaging on the target, and the imaging result s kr (f s ,f τ ) is obtained, where f s is the fast time dimension frequency, f τ is the slow time dimension frequency, so as to estimate the azimuth size S kx and the azimuth sparsity K k of the target;

[0014] Firstly, the two-dimensional ISAR coarse imaging is normalized:

[0015]

[0016] Then, a threshold T s is set, and Then, the azimuth size of the target is calculated according to the linear relationship between the Doppler frequency and the horizontal coordinate of the target scattering point ; wherein f d is the Doppler frequency generated by the relative motion between the target scattering point and the radar, v p is the radial velocity of the scattering point, λ is the wavelength of the transmitted signal, ω is the equivalent rotation angular velocity of the target, x p is the horizontal coordinate of the scattering point; and the azimuth size of the target is expressed as:

[0017]

[0018] Then, the ISAR image is normalized as follows:

[0019]

[0020] Setting threshold T h , s" kr (f τ ) is discretely expressed as a vector s" kr , the azimuthal sparsity K k of the target is kr ; and h the number of elements in the vector s"

[0021] Let the reference size be S x_ref , and the reference azimuthal resolution be p ref , then according to the azimuthal size of the target, the azimuthal resolution required by the target is:

[0022]

[0023] Further, in the first step, the threshold T s is 0.5, and the threshold T h is: 0.2.

[0024] Further, the second step is specifically:

[0025] According to the geographical positions of the radar and the target, the radar line-of-sight direction vector is determined where P k and P r are the geographical positions of the target and the radar, respectively.

[0026] The target motion speed v k is decomposed along the radar line-of-sight direction and its orthogonal direction to obtain the target radial speed v p and the tangential speed v t , wherein the tangential speed is expressed as:

[0027]

[0028] For the three-dimensional space where the target is located, a three-dimensional rectangular coordinate system (O, X, Y, Z) is established with the center of the target O as the origin. For the imaging plane, the W axis is established along the radar line-of-sight direction, and the U axis is established along the target tangential speed direction, and the unit vectors of the two axes are u u = v t / |v t |, then the radar imaging plane is represented as (O, U, W);

[0029] Next, the projection relationship is derived. Under the far-field condition, it is considered that the target ISAR image is the projection of the three-dimensional structure of the target on the imaging plane. P is a point outside the imaging plane, and Q is the projection point of point P on the imaging plane, then is perpendicular to the two unit vectors in the plane, and is expressed as:

[0030]

[0031] Let the three-dimensional coordinates of point P be (x, y, z) T , and the two-dimensional coordinates of point Q be (u, w) T , then is expressed as Express equation (6) in the form of matrix multiplication:

[0032]

[0033] The projection relationship between the three-dimensional coordinates of the scattering point and the two-dimensional coordinates on the imaging plane is constructed:

[0034]

[0035] wherein, is the projection matrix of the three-dimensional coordinates of a point in space to the imaging plane.

[0036] Further, the third step is specifically:

[0037] Design a three-dimensional Boolean matrix X with a size of N x ·N y ·N z as an optimization variable, wherein the matrix value of 1 represents the presence of a target scattering point at the point, and the strong scattering points in the two-dimensional ISAR image of the target formed by the i-th radar are extracted according to the peak value of the ISAR image intensity and represented as G i , and the projection image of the target on the imaging plane is represented as F i (X), wherein the projection function F i (·) is obtained according to the projection matrix φ of the imaging plane; the optimization function is constructed with the objective of minimizing the difference between the target two-dimensional ISAR image and the projection image:

[0038]

[0039] wherein, N R indicates that a total of N R radars image the target, and λ is a regularization parameter; the reconstruction of the three-dimensional scattering point structure of the target can be realized by solving the optimization model;

[0040] Then analyze the influence of the radar node distribution on the target imaging quality;

[0041] A single two-dimensional image cannot obtain the three-dimensional structure of the target, and at least two radars image each target to realize the positioning of the three-dimensional scattering point of the target; let two points P1 and P2 in space have two radars imaging the target, and the projection points of P1 and P2 on the two imaging planes are Q i1 and Qi2 , Q j1 , Q j2 , set The normal vector of the imaging plane i is n i , θ i is the angle between n and n i ; the distance between Q i1 and Q i2 is expressed as:

[0042]

[0043] When l i ≥ ρ i , P1 and P2 can be distinguished on the imaging plane i, and ρ i is the two-dimensional resolution of the imaging plane i, and its equivalent condition is: If two points can be distinguished on at least one imaging plane, correct reconstruction of the two points can be achieved;

[0044] Therefore, the condition for correct reconstruction of the target is:

[0045]

[0046] C0 is the set of target scattering points, and N is the number of radars observing the target; the condition is converted into a condition related only to the distribution of radar nodes, which is expressed as:

[0047]

[0048] where θ ij is the angle between the imaging planes of the two radars, ρ rec is the three-dimensional resolution of the target, and θ ij is obtained by the angle between the normal vectors of the two imaging planes;

[0049] The cross product of the unit vector bases of the imaging planes gives the normal vectors n i , n j of the i-th plane and the j-th plane, and the angle between the imaging plane normal vectors is the angle between the imaging planes:

[0050]

[0051] If the two observation radars satisfy the above condition, the target scattering points can be distinguished on at least one imaging plane of one of the radars, so correct reconstruction of the target can be achieved, and further analysis of the influence of the number of observation radars on the imaging quality is performed.

[0052] Further, the further analysis of the influence of the number of observation radars on the imaging quality in the third step is specifically:

[0053] Let the target correct reconstruction result and some wrong reconstruction result be X and X', bring the reconstruction result into the loss function in the optimization model of formula (9), in order to ensure that the final solution result of the model is the target correct reconstruction result, the following conditions need to be met:

[0054]

[0055] Let X contain two scattering points P1 and P2, and X' contain only one scattering point, the plane that can correctly distinguish P1 and P2 has N P , then:

[0056]

[0057] The first half of the formula is negative, it can be seen that the greater N P , the greater the probability that the above inequality is established; for most scattering point pairs, the more the number of radars observing the target, the more the number of imaging planes N P that can correctly distinguish them; therefore, to this extent, we think that the more the number of radars observing the target, the lower the probability of the target being reconstructed incorrectly, and the higher the target imaging quality.

[0058] Further, the fourth step is specifically:

[0059] A multi-target three-dimensional imaging resource scheduling model of networked radars is constructed, radars are allocated to each target for three-dimensional imaging, and the consumption of networked radar resources is minimized under the premise of ensuring a certain imaging quality;

[0060] The total time of scheduling is the shortest as our objective function, the constraint condition of the target imaging requirement is set, there are M radars and N targets to be imaged, define an M*N dimensional Boolean matrix W as an optimization variable to represent the scheduling scheme, W(m,n) = 1 represents that the mth radar observes n targets; let the total time used by the networked radars to complete multi-target three-dimensional imaging be T, which is solved by the following steps:

[0061] Step 4.1 extract from the scheme W which targets each radar needs to observe, and record in the target set TSet(m) of the radar;

[0062] Step 4.2 according to the target flight speed, the distance between the target and the radar, and the yaw angle, the target priority p k is comprehensively evaluated, and then the imaging required azimuth resolution is calculated to obtain the imaging required coherent accumulation time T k and observation dimension M k ;

[0063] Step 4.3 Each radar assigns sparse pulses to targets according to priority in turn until all targets in the target set are imaged;

[0064] Step 4.4 Calculate the time T consumed by each radar to complete the two-dimensional ISAR imaging of all targets in its target set m Then select the longest time as the total time T used by the networked radars to complete multi-target three-dimensional imaging;

[0065] The above steps are defined as Tconsum(W), and the objective function is represented as:

[0066] min Tconsum(W) (16)

[0067] Then determine whether the scheme W satisfies the target reconstruction condition;

[0068] Suppose the radar observes all targets, then calculate the angle between the normal vectors of the imaging planes of each pair of radars observing N targets, and calculate the minimum angle that the imaging planes of the two radars must satisfy according to formula (12). Record them in The matrix θ RT and θ pho , representing a total of radar node combination ways;

[0069] Then perform a bitwise AND operation between the elements of each column in W according to the order of θ RT and θ ρ , to obtain the matrix RT of dimension, RT(k,n) = 1 indicates that target n is observed by two radars in the kth radar combination;

[0070] Then determine whether the scheme W satisfies the reconstruction condition of each target:

[0071] Satis = [(RT ⊙ θ RT ) > θ pho ] (17)

[0072] If there is a radar combination in Satis that satisfies the condition in any column, i.e. Then consider that the scheme W can achieve correct reconstruction of each target; at the same time, in order to reduce the probability of error reconstruction, each target is observed by at least 3 radars;

[0073] List the resource scheduling model:

[0074]

[0075] Further, the fourth step sets the constraint condition of the requirement of target imaging as follows: 1. Each target has at least two radars that can meet the reconstruction condition of any scattering point, that is, there are two radars that meet the condition of formula (12); 2. The number of radars observing the target is as large as possible to reduce the probability of false reconstruction, and each target has at least three radars for imaging; 3. Each radar can use a sparse ISAR imaging algorithm based on compressed sensing to observe multiple targets, so the pulse sequence allocated to each target needs to meet the requirement of sparse reconstruction.

[0076] Further, the fifth step is as follows:

[0077] The A2C method deep reinforcement learning method is used to solve the resource scheduling model, and the key elements of reinforcement learning are designed.

[0078] Step 5.1 State: Each radar node allocation scheme is regarded as a State, and the radar node selection scheme of each target is set as S n , then the State of all target radar node selection schemes is S = (S1, S2,..., Sm) n , where S n is a Boolean vector with a length of m;

[0079] Step 5.2 Action: the action is set as the target number n and the radar node selection scheme S k : a = (n, S k ), that is, each action only changes the radar node selection scheme of one target, wherein some radar node selection schemes that do not meet the condition are filtered out according to the constraint condition of the resource scheduling model, and the action space is pruned;

[0080] Step 5.3 State Transition: directly change the radar node selection scheme of target n to S k according to the action a = (n, S k );

[0081] Step 5.4 State Value: the state value is set to be consistent with the objective function of the scheduling model, and is set as the negative value of the total time of the networking radars completing the three-dimensional imaging task of multiple targets;

[0082] Step 5.5 Reward: the reward of each action is set as the state value of the next state minus the state value of the current state, that is:

[0083] Reward = SV (S n )- SV (S c) (19)

[0084] where S n denotes the next state, S c denotes the current state, and SV(·) denotes the state value; when the discount factor γ is set to a number tending to 1, the cumulative return return of each round of training is approximately equal to the state value of the last state of the round of training minus the state value of the initial state, i.e.

[0085]

[0086] where r is the reward value, and as can be seen from equation (20), the state value of the initial state SV(S0) is constant, and the state with the maximum state value is obtained by using the method of reinforcement learning.

[0087] Further, the sixth step specifically includes:

[0088] The specific steps of solving the resource scheduling model by using the A2C method include:

[0089] Step 6.1 initializing the policy network parameter Θ, the value network parameter Ω, and best_value = -10000

[0090] Step 6.2 for training round from 1 to 500:

[0091] Step 6.3 done = False

[0092] Step 6.4 while not done:

[0093] Step 6.5 obtaining the action a c according to the policy π(Θ) and the current state S

[0094] Step 6.6 obtaining the next state S n and the reward r after performing the action n

[0095] Step 6.7 calculating the TD error and updating the two network parameters Θ and Ω

[0096] Step 6.8 if SV(S n ) > best_value

[0097] Step 6.9 best_value = SV(S n )

[0098] Step 6.10 end if

[0099] Step 6.11 if SV(Sn )==best_value

[0100] Step 6.12 done=True

[0101] Step 6.13 end if

[0102] Step 6.14 S c =S n

[0103] Step 6.15 end while

[0104] Step 6.16 end for

[0105] After the algorithm converges, the optimal radar node scheduling scheme can be obtained, then each radar assigns pulses to the targets in its target set according to the scheduling scheme according to the priority, and finally reconstructs the three-dimensional image of each target according to the three-dimensional imaging algorithm.

[0106] The present application has the following advantages: 1. The influence of radar node distribution on target three-dimensional imaging quality is considered, a networking radar multi-target three-dimensional imaging resource scheduling model is constructed, and the time resource consumption is minimized under the premise of ensuring the imaging quality. 2. The method of deep reinforcement learning is used to solve the model, and the optimal solution of the problem can be obtained. BRIEF DESCRIPTION OF DRAWINGS

[0107] Figure 1 A radar imaging plane coordinate system construction schematic diagram is shown;

[0108] Figure 2 A target scattering point projection schematic diagram is shown;

[0109] Figure 3 A target two scattering point projection schematic diagram is shown;

[0110] Figure 4 A networking radar multi-target three-dimensional imaging resource scheduling method flow chart based on deep reinforcement learning is shown;

[0111] Figure 5 A radar, target space geographical position distribution diagram is shown;

[0112] Figure 6 A target motion direction diagram is shown;

[0113] Figure 7 A reference scattering point model diagram is shown;

[0114] Fig. 8(a) shows a scatter point model diagram of target 1, Fig. 8(b) shows a scatter point model diagram of target 2, Fig. 8(c) shows a scatter point model diagram of target 3, Fig. 8(d) shows a scatter point model diagram of target 4, Fig. 8(e) shows a scatter point model diagram of target 5, and Fig. 8(f) shows a scatter point model diagram of target 6;

[0115] Figure 9 A decision network and a value network structure schematic diagram are shown;

[0116] Fig. 10(a) shows the value of the reward of each round of training of the reinforcement learning, and Fig. 10(b) shows the state value of the final state of each round of training of the reinforcement learning;

[0117] Figure 11 Three two-dimensional ISAR images of target 1 are shown;

[0118] Fig. 12(a) shows a three-dimensional reconstruction image of target 1, Fig. 12(b) shows an X-Y plane projection image of the three-dimensional reconstruction structure of target 1, Fig. 12(c) shows an X-Z plane projection image of the three-dimensional reconstruction structure of target 1, and Fig. 12(d) shows a Y-Z plane projection image of the three-dimensional reconstruction structure of target 1. DETAILED DESCRIPTION

[0119] The application is further described below in combination with embodiments and drawings.

[0120] A networking radar multi-target three-dimensional imaging resource scheduling method based on deep reinforcement learning considers the influence of radar node distribution on target three-dimensional imaging quality, analyzes the demand of target three-dimensional imaging on radar resources, and allocates radar nodes for each target, so as to minimize the consumption of radar resources under the premise of ensuring imaging quality, and specifically includes the following steps:

[0121] First step: sensing the target features and calculating the required azimuth resolution of the target;

[0122] Second step: determining the two-dimensional imaging plane of each radar on the target and the projection relationship;

[0123] Third step: analyzing the influence of radar node distribution on target three-dimensional imaging quality;

[0124] Fourth step: constructing a networking radar multi-target three-dimensional imaging resource scheduling model;

[0125] Fifth step: designing key elements of reinforcement learning;

[0126] Sixth step: solving the networking radar multi-target three-dimensional imaging resource scheduling model, and performing three-dimensional imaging on the target according to the scheduling result.

[0127] Further, the first step specifically includes:

[0128] The geographical position, the motion speed, the distance and the yaw angle between each radar and each target are obtained by tracking the targets in the detection area through the radar. kr (f s ,f τ ), wherein f s is the fast time dimension frequency, and f τ is the slow time dimension frequency, so as to estimate the azimuth size S kx and the azimuth sparsity K k of the target.

[0129] Firstly, the ISAR image is normalized as follows:

[0130]

[0131] Then, a threshold T s is set, T s is 0.5, and s The azimuth size of the target is calculated according to the linear relationship between the Doppler frequency and the horizontal coordinate of the target scattering point , wherein f d is the Doppler frequency generated by the relative motion between the target scattering point and the radar, v p is the radial velocity of the scattering point, λ is the wavelength of the transmitted signal, ω is the equivalent rotation angular velocity of the target, x p is the horizontal coordinate of the scattering point; and the azimuth size of the target is expressed as:

[0132]

[0133] Then, the ISAR image is normalized as follows:

[0134]

[0135] A threshold T h is set, T h is 0.2, and s kr (f τ ) is discretely expressed as a vector s kr , wherein the azimuth sparsity K k of the target is the number of elements greater than T h in the vector s kr .

[0136] The reference size is S x_ref , and the reference azimuth resolution is ρ refThen the azimuth resolution required by the target according to the target azimuth dimension is:

[0137]

[0138] Further, the second step is specifically:

[0139] First, the radar line of sight (LOS) vector is determined according to the geographical positions of the radar and the target Wherein, P k and P r are the geographical positions of the target and the radar respectively;

[0140] Then, the target motion speed v k is decomposed along the radar LOS direction and the orthogonal direction thereof to obtain the target radial speed v p and the tangential speed v t , wherein the tangential speed is expressed as:

[0141]

[0142] For the three-dimensional space where the target is located, a three-dimensional rectangular coordinate system (O, X, Y, Z) with the center of the target O as the origin is established; for the imaging plane, as shown in Figure 1 , the W axis is established along the radar LOS direction, and the U axis is established along the target tangential speed direction, and the unit vectors of the two axes are respectively u u and v t / |v t |, and the radar imaging plane is represented as (O, U, W);

[0143] Next, the projection relationship is derived. Under the far-field condition, it is considered that the ISAR image of the target is the projection of the three-dimensional structure of the target on the imaging plane; as shown in Figure 2 , P is a point outside the imaging plane, and Q is the projection point of point P on the imaging plane, then is perpendicular to the two unit vectors in the plane, and is expressed as:

[0144]

[0145] Suppose the three-dimensional coordinates of point P are (x, y, z) T , and the two-dimensional coordinates of point Q are (u, w) T , then is expressed as The formula (6) is expressed in the form of matrix multiplication:

[0146]

[0147] The projection relationship between the three-dimensional coordinates of the scattering points and the two-dimensional coordinates on the imaging plane is constructed:

[0148]

[0149] wherein, is the projection matrix of the three-dimensional coordinates of a point in space to the imaging plane.

[0150] Further, the third step is specifically:

[0151] The application adopts a three-dimensional imaging method based on imaging projection to image the target, and the main idea is as follows: multiple radars are selected to image the target at different stations and different angles, and then the three-dimensional scattering point structure of the target is reconstructed according to the obtained multiple two-dimensional ISAR images and the corresponding projection relationship; according to the idea, a three-dimensional Boolean matrix X with a size of N x * N y * N z is designed as an optimization variable, wherein the matrix value of 1 represents that there is a target scattering point at the point. The strong scattering points in the two-dimensional ISAR image of the target imaged by the i-th radar are extracted according to the peak value of the ISAR image intensity, and are represented as G i , and the projection image of the target on the imaging plane is represented as F i (X), wherein the projection function F i (·) can be obtained according to the projection matrix φ of the imaging plane. An optimization function is constructed by taking the minimization of the difference between the target two-dimensional ISAR image and the projection image as the target:

[0152]

[0153] wherein N R indicates that a total of N R radars image the target, and λ is a regularization parameter. The three-dimensional scattering point structure of the target can be reconstructed by solving the optimization model.

[0154] Then the influence of the radar node distribution on the target imaging quality is analyzed.

[0155] A single two-dimensional image cannot obtain the three-dimensional structure of the target, and at least two radars image each target to realize the positioning of the three-dimensional scattering points of the target. As shown in Figure 3 , two radars image the target, and the projection points of P1 and P2 on the two imaging planes are Q i1 , Q i2 , Q j1 , and Q j2 , respectively, wherein the normal vector of the imaging plane i is n i , and θ ifor With n i The angle between Q and Q. i1 With Q i2 The distance between them is expressed as:

[0156]

[0157] When l i ≥ρ i At that time, P1 and P2 can be distinguished on the imaging plane i, ρ i For the two-dimensional resolution of imaging plane i, the equivalent condition is: If two points can be distinguished at least on an imaging plane, then the correct reconstruction of the two points can be achieved.

[0158] Therefore, the conditions for achieving a correct reconstruction of the target are expressed as follows:

[0159]

[0160] C0 represents the set of target scattering points, and N represents the number of radars observing the target. This condition is transformed to depend only on the distribution of radar nodes, as follows:

[0161]

[0162] Where θ ij ρ is the angle between the two radar imaging planes. rec For the target 3D resolution, θ ij Obtained by the angle between the normal vectors of the two imaging planes:

[0163] The cross product of the unit vector basis of the imaging plane yields the normal vector n of the i-th and j-th planes. i n j The angle between the normal vectors of the imaging plane is the angle between the imaging planes:

[0164]

[0165] If the two observation radars meet the above conditions, the target scattering point pair can be distinguished on the imaging plane of at least one of the radars, thus enabling the correct reconstruction of the target.

[0166] Furthermore, the impact of the number of observation radars on imaging quality is analyzed.

[0167] Let the correct reconstruction result and the incorrect reconstruction result be X and X', respectively. Substitute the reconstruction result into the loss function in the optimization model of equation (9). In order to ensure that the final solution of the model is the correct reconstruction result, the following conditions need to be met:

[0168]

[0169] Suppose X contains two scattering points P1, P2, and X' contains only one scattering point, the planes that can correctly distinguish P1 and P2 are N P , then:

[0170]

[0171] The first half of the formula is negative, so it can be seen that the greater N P , the greater the probability that the above inequality holds. For most scattering point pairs, the more radars that image the target, the more imaging planes N P that can correctly distinguish them. Therefore, to this extent, we believe that the more radars that observe a single target, the lower the probability of the target being incorrectly reconstructed, and the higher the quality of the target image.

[0172] Further, the fourth step is specifically:

[0173] According to the above analysis, a multi-target three-dimensional imaging resource scheduling model of networked radars is constructed, and radars are allocated to each target for three-dimensional imaging, so that the networked radar resource consumption is minimized under the premise of ensuring a certain imaging quality.

[0174] First of all, the objectives of the problem are three:

[0175] 1. Achieve three-dimensional imaging of all targets;

[0176] 2. Ensure that each target meets certain imaging quality requirements;

[0177] 3. The total time used by the networked radars for multi-target three-dimensional imaging is the shortest.

[0178] In order to avoid solving complex multi-objective optimization problems, in the resource scheduling model of the present application, we take the shortest total time as our objective function, and the requirements for target imaging are embodied in the constraint conditions. The constraint conditions are set as follows:

[0179] 1. Each target has at least two radars that can meet the reconstruction conditions for any scattering point, i.e., there are two radars that meet the condition of formula (12);

[0180] 2. The number of radars observing the target is as large as possible to reduce the probability of incorrect reconstruction, and each target has at least three radars imaging it;

[0181] 3. Each radar can use a sparse ISAR imaging algorithm based on compressed sensing to observe multiple targets simultaneously, so the pulse sequence allocated to each target needs to meet the sparse reconstruction requirements.

[0182] Given M radars and N targets to be imaged, we define an M*N dimensional Boolean matrix W as an optimization variable to represent the scheduling scheme, where W(m,n) = 1 represents the m-th radar imaging the n targets. Let T be the total time T taken for the networked radars to complete multi-target 3D imaging, which can be solved using the following steps:

[0183] Step 4.1 Extract from scheme W which targets each radar needs to observe and record them in the radar target set TSet(m);

[0184] Step 4.2 Prioritize targets based on their flight speed, distance from radar, and yaw angle. k A comprehensive evaluation is conducted, and then the coherent accumulation time T required for imaging is calculated based on the azimuth resolution required for target imaging. k and observation dimension M k ;

[0185] Step 4.3 Each radar assigns sparse pulses to the targets in sequence according to priority, until imaging of all targets in the target set is completed;

[0186] Step 4.4 Calculate the time T taken by each radar to complete two-dimensional ISAR imaging of all targets in its target set. m Then, the longest time is selected as the total time T used by the networked radar to complete multi-target three-dimensional imaging.

[0187] If we define the above steps as Tconsum(W), then the objective function is expressed as:

[0188] min Tconsum(W)(16)

[0189] Then determine whether scheme W satisfies the target reconstruction conditions.

[0190] Assuming the radar observes all targets, the angle between the normal vectors of the imaging planes of every two radars observing N targets can be calculated. Then, according to equation (12), the minimum angle between the imaging planes of these two radars that must satisfy the reconstruction condition can be calculated, and these angles are recorded. A 2D matrix θ RT and θ pho middle, The total number of representatives is Various radar node combination methods.

[0191] Then, for each column of W, elements are processed according to θ. RT and θ ρ Perform a pairwise bitwise AND operation on each element in sequence to obtain... The matrix RT, RT(k, n) = 1 indicates that the target n is observed by two radars in the kth radar combination. Then, it is judged whether the scheme W satisfies the reconstruction condition of each target:

[0192] Satis = [(RT o &theta RT ) > &theta pho ] (17)

[0193] If any column in Satis has a radar combination satisfying the condition, that is, then it is considered that the scheme W can realize correct reconstruction of each target. Meanwhile, in order to reduce the probability of error reconstruction, it is hoped that each target is observed by at least three radars.

[0194] Based on the above, the resource scheduling model can be listed as follows:

[0195]

[0196] Further, the fifth step is specifically as follows:

[0197] The advantage actor-critic (A2C) deep reinforcement learning method is used to solve the resource scheduling model, and the key elements of reinforcement learning are designed as follows:

[0198] Step 5.1 State: each radar node allocation scheme is regarded as a State, and the radar node selection scheme of each target is set as S n , then the State of all target radar node selection schemes is S = (S1, S2,..., Sm) n ), wherein S n is a Boolean vector with a length of m, for example, S1 = (0, 1, 1, 0) represents that there are a total of 4 radars, and target 1 is observed by radars 2 and 3;

[0199] Step 5.2 Action: the action is set as the target sequence number n + radar node selection scheme S k : a = (n, S k ), that is, each action only changes the radar node selection scheme of one target, wherein some radar node selection schemes that do not meet the condition are filtered out according to the constraint condition of the resource scheduling model, and the action space is pruned;

[0200] Step 5.3 State Transition: the radar node selection scheme of target n is changed to S k according to the action a = (n, S k );

[0201] Step 5.4 State Value: The state value is set to be consistent with the objective function of the scheduling model, which is the negative value of the total time for the multi-target three-dimensional imaging task of the networked radar;

[0202] Step 5.5 Reward: In reinforcement learning, reward is used to convey the goal you want to achieve. Since the goal of reinforcement learning is to maximize cumulative return, and the resource scheduling model we set is to find the state with the maximum state value, we set the reward of each action as the state value of the next state minus the state value of the current state, that is:

[0203] Reward=SV(S n )-SV(S c ) (19)

[0204] Where S n represents the next state, S c represents the current state, and SV(·) represents the state value. When the discount factor γ is set to a number close to 1, the cumulative return return of each round of training is approximately equal to the state value of the last state of the round of training minus the state value of the initial state, that is:

[0205]

[0206] Where r is the reward value. As can be seen from the formula, the state value of the initial state SV(S0) is constant, and the method of using reinforcement learning can obtain the state with the maximum state value.

[0207] Further, the sixth step is specifically:

[0208] The present application uses the A2C deep reinforcement learning method to solve the resource scheduling model. Actor-Critic is a classic policy-based deep reinforcement learning method, which consists of two networks, Actor and Critic. The Actor is a decision-making network that can make decisions and actions based on input states, while the Critic is a value network that evaluates the goodness of the selected actions. The Actor-Critic method continuously interacts with the environment to obtain feedback, and then uses the feedback information to train the parameters of the two networks, and finally realizes the convergence of the network. The A2C method further reduces the estimation variance of the action value by introducing a baseline based on the Actor-Critic method.

[0209] Since the resource scheduling problem is an optimization problem, we need to first convert the problem into a Markov decision process. In action design, we design each action to change only one target radar node selection scheme, so we will perform several actions from an initial state before obtaining the optimal solution. Since the optimal solution is unknown, we set the end point of each training dynamically, and the best solution obtained is used as the end point of this round of training. When we get the current optimal solution or find a solution with higher state value, we end this round of training. In summary, we convert the problem into a Markov decision process, and then use the A2C method to solve the model.

[0210] The specific steps of solving the resource scheduling model using the A2C method are as follows:

[0211]

[0212]

[0213] After the algorithm converges, the optimal radar node scheduling scheme is obtained, and then each radar assigns pulses to the targets in its target set according to the priority based on the scheduling scheme, and finally reconstructs the three-dimensional image of each target based on the three-dimensional imaging algorithm.

[0214] Simulation example

[0215] Figure 4 The flowchart of the multi-target three-dimensional imaging resource scheduling method for networked radars based on deep reinforcement learning is shown. The algorithm first perceives the target features, determines the two-dimensional imaging plane of each radar for the target and the projection relationship, and then analyzes the impact of radar node distribution on the quality of target three-dimensional imaging. On this basis, a multi-target three-dimensional imaging resource scheduling model for networked radars is constructed, the key elements of reinforcement learning are designed, and the A2C deep reinforcement learning method is used to solve the resource scheduling model. Finally, each radar performs two-dimensional ISAR imaging on the target based on the scheduling result, and uses the three-dimensional imaging algorithm based on imaging projection to realize target three-dimensional scatter point reconstruction based on multiple two-dimensional ISAR images and their corresponding projection relationships.

[0216] Suppose the radar transmits a linear frequency modulation signal, the transmit signal carrier frequency is f c = 10 GHz, the bandwidth is B = 300 MHz, the range resolution is p r = 0.5 m, the pulse repetition frequency PRF = 1000 Hz, the pulse width T p = 1 s, the target reference azimuth size S x_ref = 25 m, the required azimuth resolution is p ref = 0.5 m, and the target three-dimensional resolution is p rec = 1 m.

[0217] A networked radar consisting of 5 radars is set up to observe 6 targets in the observation range, and their spatial geographic positions and target motion directions are shown in Figure 5 、 Figure 6 The spatial position coordinates of the radars and targets and the target velocity vectors are shown in Table 2 and Table 3, respectively.

[0218] Table 2 Radar and target spatial position coordinate table

[0219]

[0220] Table 3 Target velocity vector table

[0221]

[0222] The three-dimensional scattering point model of the target is unified with the scattering point model shown in Figure 7 , and the corresponding rotation is performed around the center of the scattering point model according to the target motion direction. The scattering point models of the 6 targets are shown in Figure 8(a)-Figure 8(f) .

[0223] The networked radar multi-target three-dimensional imaging resource scheduling method based on deep reinforcement learning specifically includes the following steps:

[0224] Step 1: Perceive the target features and calculate the required azimuth resolution of the target

[0225] Through the tracking results of each target in the detection area by the radars, the geographic position, motion speed, and distance and yaw angle between each target and each radar are obtained; then each radar transmits a small amount of pulses to perform two-dimensional ISAR rough imaging on the target, and through the analysis of the image, the azimuth size S kx and the azimuth sparsity K k of the target observed by each radar can be estimated, and the results are shown in Table 4 and Table 5, respectively.

[0226] Table 4 Azimuth size of target (m)

[0227]

[0228]

[0229] Table 5 Target sparsity

[0230] Target 1 Target 2 Target 3 Target 4 Target 5 Target 6 Radar 1 45 41 32 54 53 33 Radar 2 46 23 43 46 42 37 Radar 3 44 35 48 45 52 47 Radar 4 33 44 24 42 36 38 Radar 5 48 38 41 45 42 38

[0231] Step 2: Determine the two-dimensional imaging plane of each radar to the target and the projection relationship

[0232] First, the radar line of sight (LOS, line of sight) vector is determined according to the geographic position of the radar and the target Then the target motion velocity v k Decomposing along the radar LOS direction and its orthogonal direction, the target radial velocity v p and the tangential velocity v t Where the tangential velocity is expressed as:

[0233]

[0234] The W-axis is established along the radar LOS direction, and the U-axis is established along the target tangential velocity direction, and the unit vectors of the two axes are respectively u u =v t / |v t |, then the radar imaging plane can be represented as (O, U, W).

[0235] Under the far-field condition, the target ISAR image can be considered as the projection of the three-dimensional structure of the target on the imaging plane. Then the projection relationship between the three-dimensional coordinates of the scattering point and the two-dimensional coordinates on the imaging plane can be determined as:

[0236]

[0237] Where, is the projection matrix of a point in space to the imaging plane.

[0238] Third step: analyze the influence of radar node distribution on the quality of target three-dimensional imaging

[0239] The three-dimensional imaging method based on imaging projection is adopted to image the target, and the main idea is: selecting multiple radars to perform two-dimensional ISAR imaging on the target at different stations and different angles, and then reconstructing the three-dimensional scattering point structure of the target according to the obtained multiple two-dimensional ISAR images and the corresponding projection relationship. According to the idea, a three-dimensional Boolean matrix X with a size of N x ·N y ·N z is designed as an optimization variable, wherein the matrix value of 1 represents that there is a target scattering point at the point. According to the peak value of the ISAR image intensity, the strong scattering points in the two-dimensional ISAR image of the target formed by the i-th radar are extracted and represented as G i The projection image of the target on the imaging plane is represented as F i (X), wherein the projection function F i (·) can be obtained according to the projection matrix φ of the imaging plane. The optimization function is constructed by taking the minimization of the difference between the target two-dimensional ISAR image and the projection image as the target:

[0240]

[0241] Where, N RThis indicates that there are a total of N. R The radar images the target, where λ is the regularization parameter. By solving this optimization model, the three-dimensional scattering point structure of the target can be reconstructed.

[0242] Then, the impact of radar node distribution on the three-dimensional imaging quality of the target is analyzed.

[0243] A single two-dimensional image cannot reveal the three-dimensional structure of a target. To locate the three-dimensional scattering points of a target, each target must be imaged by at least two radars. Let there be two points P1 and P2 in space, such as... Figure 3 As shown, two radars image the target, and the projection points of P1 and P2 on the two imaging planes are Q1 and Q2, respectively. i1 Q i2 Q j1 Q j2 ,set up The normal vector of the imaging plane i is n i θ i for With n i The angle between Q and Q. i1 With Q i2 The distance between them can be expressed as:

[0244]

[0245] When l i ≥ρ i At that time, P1 and P2 can be distinguished on the imaging plane i, ρ i For the two-dimensional resolution of imaging plane i, the equivalent condition is: If two points can be distinguished at least on an imaging plane, then the correct reconstruction of the two points can be achieved.

[0246] Therefore, the conditions for achieving a correct reconstruction of the target can be expressed as:

[0247]

[0248] C0 represents the set of target scattering points, and N represents the number of radars observing the target. This condition, which depends only on the distribution of radar nodes, can be expressed as:

[0249]

[0250] Where θ ij ρ is the angle between the two radar imaging planes. rec θ represents the target 3D resolution. ij It can be obtained from the angle between the normal vectors of the two imaging planes:

[0251] The cross product of the unit vector basis of the imaging plane gives the normal vector n of the i-th plane and the j-th plane i j The angle between the imaging plane normal vectors is the angle between the imaging planes:

[0252]

[0253] If the two observation radars satisfy the above conditions, the target scattering points can be distinguished on at least one of the imaging planes of the radars, so that the correct reconstruction of the target can be achieved.

[0254] Further, we analyze the influence of the number of observation radars on the imaging quality.

[0255] Let the correct reconstruction result of the target and a certain incorrect reconstruction result be X and X', and bring the reconstruction result into the loss function in the optimization model of formula (3). In order to ensure that the final solution of the model is the correct reconstruction result of the target, the following conditions need to be met:

[0256]

[0257] Suppose that X contains two scattering points P1 and P2, and X' contains only one scattering point. The planes that can correctly distinguish the two scattering points P1 and P2 have N P , then:

[0258] The first half of the formula is negative, so it can be seen that the larger N P , the greater the probability that the above inequality holds. For most scattering point pairs, the more radars that image the target, the more imaging planes N P that can correctly distinguish them. Therefore, to this extent, we believe that the more radars that observe a single target, the lower the probability of incorrect reconstruction of the target, and the higher the quality of the target imaging.

[0259] Fourth step: constructing a multi-target three-dimensional imaging resource scheduling model for networked radars

[0260] A multi-target three-dimensional imaging resource scheduling model for networked radars is constructed, and radars are allocated to each target for three-dimensional imaging. Under the premise of ensuring a certain imaging quality, the consumption of networked radar resources is minimized.

[0261] In the embodiment, there are 5 radars and 6 targets to be imaged. A 5*6-dimensional Boolean matrix W is defined as an optimization variable to represent the scheduling scheme, and W(m, n) = 1 represents that the m-th radar images the n targets. Let the total time used by the networked radars to complete multi-target three-dimensional imaging be T, which can be solved by the following steps:

[0262] ​1. Extract from scheme W which targets each radar needs to observe and record them in the radar target set TSet(m);

[0263] 2. Prioritize targets based on their flight speed, the distance between the target and the radar, and their yaw angle. k A comprehensive evaluation is conducted, and then the coherent accumulation time T required for imaging is calculated based on the azimuth resolution required for target imaging. k and observation dimension M k ;

[0264] 3. Each radar assigns sparse pulses to the targets in sequence according to priority, until imaging of all targets in the target set is completed;

[0265] 4. Calculate the time T taken by each radar to complete two-dimensional ISAR imaging of all targets in its target set. m Then, the longest time is selected as the total time T used by the networked radar to complete multi-target three-dimensional imaging.

[0266] If we define the above steps as Tconsum(W), then the objective function can be expressed as:

[0267] min Tconsum(W)(10)

[0268] Assuming the radar can observe all targets, the angle between the normal vectors of the imaging planes of every two radars observing the six targets can be calculated. Then, according to equation (12), the minimum angle between the imaging planes of these two radars that must satisfy the reconstruction condition can be calculated, and these angles are recorded. A 2D matrix θ RT and θ ρ middle, The total number of representatives is Various radar node combination methods.

[0269] Then, for each column of W, elements are processed according to θ. RT and θ pho Perform a pairwise bitwise AND operation on each element in sequence to obtain... A matrix RT is given, where RT(k,n) = 1 indicates that target n is observed by two radars in the k-th radar combination. Then, it is determined whether scheme W satisfies the reconstruction conditions for each target:

[0270] Satis=[(RT⊙θ RT )>θ pho (11)

[0271] If any column in Satis contains radar combinations that satisfy the conditions, i.e. Therefore, it is considered that the scheme W can realize the correct reconstruction of each target.

[0272] In summary, the resource scheduling model can be listed as follows:

[0273]

[0274] Fifth step: design the key elements of reinforcement learning

[0275] The Advantage actor-critic (A2C) deep reinforcement learning method is adopted to solve the resource scheduling model, and the key elements of reinforcement learning are designed as follows:

[0276] 1. State: each radar node allocation scheme is regarded as a state. n The state of all target radar node selection schemes is S=(S1, S2,..., Sm). n ), wherein S n is a Boolean vector with a length of m, for example, S1=(0, 1, 1, 0) represents that there are a total of 4 radars, and target 1 is observed by radars 2 and 3.

[0277] 2. Action: the action is set as the target sequence number n and the radar node selection scheme S k : a=(n, S k ), that is, each action only changes the radar node selection scheme of one target.

[0278] 3. State transition: the radar node selection scheme of target n is changed to S k according to the action a=(n, S k ).

[0279] 4. State value: the state value is consistent with the objective function of the scheduling model, and is set as the negative value of the total time of the networking radars to complete the multi-target three-dimensional imaging task.

[0280] 5. Reward: the reward of each action is set as the state value of the next state minus the state value of the current state, that is:

[0281] Reward=SV(S n )-SV(S c ) (13)

[0282] wherein S n denotes the next state, S c denotes the current state, and SV(·) denotes the state value. When the discount factor γ is set to a number tending to 1, the cumulative return return of each round of training can be approximated to the state value of the last state of the round of training minus the state value of the initial state, that is:

[0283]

[0284] wherein r is the reward value. As can be seen from the formula, the state value of the initial state SV(S0) is constant, and the state with the maximum state value can be obtained by using the method of reinforcement learning.

[0285] Step 6: solving the multi-target three-dimensional imaging resource scheduling model of the networking radar, and performing three-dimensional imaging on the target according to the scheduling result

[0286] The method of A2C deep reinforcement learning is used to solve the resource scheduling model. The decision network and the value network use a three-layer fully connected neural network, which includes an input layer, a hidden layer and an output layer. The number of neurons in the hidden layer is 128, and ReLU is used as the activation function. The network structure is as shown in Figure 9 .

[0287] The learning rate of the decision network is set to lr_a=1*10 -3 , the learning rate of the value network is set to lr_v=1*10 -3 , the value discount rate is set to γ=0.98, and the total number of iterations is set to 500.

[0288] The network training result is shown in Figs. 10(a) and 10(b). From the figures, we can see that the algorithm finds the optimal solution after about 100 rounds of training, and then the network basically converges.

[0289] The optimal scheduling scheme obtained is shown in Table 6.

[0290] Table 6 Optimal scheduling scheme

[0291]

[0292] The total imaging time of the scheme is 2.055 seconds, and the pulse utilization rates of the radars are 0.75, 0.88, 0.73, 0.52 and 0.63, respectively.

[0293] Each radar images the target according to the scheme, acquiring multiple 2D ISAR images of the target from different perspectives. Then, based on the target 3D imaging algorithm, the 3D structure of the target is reconstructed using the multiple 2D ISAR images and their corresponding projection relationships. The mean square error of each target reconstruction result is shown in Table 7.

[0294] Table 7. Mean Squared Error (MSE) for Reconstruction of Each Target

[0295]

[0296] Among them, the three two-dimensional imaging results of target 1 are as follows: Figure 11 As shown, the three-dimensional imaging result of target 1 is as follows: Figure 12(a)-Figure 12(d) As shown in the figure, both the target image and the MSE index demonstrate that the imaging quality obtained by the algorithm proposed in this invention meets the requirements.

[0297] This invention further compares the performance of the proposed algorithm with an existing resource scheduling algorithm for multi-target 3D imaging of networked radar. Algorithm 2 constructs a resource scheduling model with the goal of minimizing the total imaging time, requiring each target to be observed by three non-collinear radars. However, it does not give much consideration to the impact of radar node distribution on imaging quality.

[0298] The scheduling scheme obtained by using Algorithm 2 to solve the scenario in this embodiment is shown in Table 8.

[0299] Table 8 Scheduling Scheme of Algorithm 2

[0300]

[0301] The total imaging time obtained by this scheme is 2.055 seconds, and the pulse utilization rates of each radar are 0.74, 0.67, 0.55, 0.52, and 0.41, respectively, which are generally lower than the radar pulse utilization rate of the algorithm proposed in this invention.

[0302] The mean squared error (MSE) of the target reconstruction obtained by the two algorithms is shown in Table 9.

[0303] Table 9. Mean Squared Error (MSE) of the two algorithms for target reconstruction

[0304]

[0305] As can be seen from the data in the table, for targets 4 and 5, the imaging effect of the proposed algorithm is significantly better than that of algorithm 2 when the scheduling results of the proposed algorithm are obtained, thus proving the effectiveness of the proposed algorithm.

[0306] The present application has the following advantages: 1. The influence of radar node distribution on target three-dimensional imaging quality is considered, a multi-target three-dimensional imaging resource scheduling model of networked radar is constructed, and the time resource consumption is minimized under the premise of ensuring the imaging quality. 2. The method of deep reinforcement learning is used to solve the model, and the optimal solution of the problem can be obtained.

Claims

1. A resource scheduling method for multi-target three-dimensional imaging in networked radar based on deep reinforcement learning, specifically including the following steps: Step 1: Perceive the target features and calculate the required azimuth resolution for the target; Step 2: Determine the two-dimensional imaging plane and projection relationship of each radar on the target; Step 3: Analyze the impact of radar node distribution on the three-dimensional imaging quality of the target; Step 4: Construct a multi-target 3D imaging resource scheduling model for networked radar; Step 5: Design the key elements of reinforcement learning; Step 5 specifically includes: The resource scheduling model is solved using the A2C deep reinforcement learning method, and the key elements of reinforcement learning are designed: Step 5.1 State: Each radar node allocation scheme is considered a State, and the radar node selection scheme for each target is S. n Then State represents the radar node selection scheme for all targets: S = (S1, S2, ..., S...). n ), where S n Let m be a Boolean vector; Step 5.2 Action: Set the action to target number n + radar node selection scheme S k : a=(n,S k This means that each action only changes the radar node selection scheme for one target. Based on the constraints of the resource scheduling model, some radar node selection schemes that do not meet the conditions will be filtered out, and the action space will be pruned. Step 5.3 State Transition: Directly based on the action a = (n, S) k Change the radar node selection scheme for target n to S k ; Step 5.4 State Value: The state value is aligned with the objective function of the scheduling model and is set as the negative of the total time for the networked radar to complete the multi-target 3D imaging task. Step 5.5 Reward: Set the reward for each action to the state value of the next state minus the state value of the current state, i.e.: Reward=SV(S n )-SV(S c ) (19) Among them, S n S represents the next state. c Let SV represent the current state, and SV(·) represent the state value. When the discount factor γ is set to a number approaching 1, the cumulative return for each training round is approximately equal to the state value of the last state in that round minus the state value of the initial state, i.e.: Where r is the reward value, it can be seen from equation (20) that the initial state value SV(S0) is constant, and the state with the maximum state value is obtained by using reinforcement learning. Step 6: Solve the resource scheduling model for multi-target 3D imaging of the networked radar, and perform 3D imaging of the targets based on the scheduling results.

2. The method for scheduling multi-target three-dimensional imaging resources of networked radar based on deep reinforcement learning as described in claim 1, wherein the first step specifically comprises: By tracking targets within the detection area using radar, the geographical location, speed, distance to each radar, and yaw angle of each target are obtained. Then, each radar emits a small number of pulses to perform coarse two-dimensional ISAR imaging of the targets, yielding the imaging result s. kr (f s ,f τ ), where f s For the fast time dimension frequency, f τ The frequency of the slow time dimension is used to estimate the target's azimuth dimension S. kx and azimuth sparsity K k ; First, the coarse 2D ISAR image is normalized: Then, set the threshold T. s ,remember Based on the linear relationship between the Doppler frequency and the x-coordinate of the target scattering point... Calculate the target azimuth dimension; where f d v is the Doppler frequency generated by the relative motion between the target scattering point and the radar. p Let λ be the radial velocity of the scattering point, λ be the wavelength of the emitted signal, ω be the equivalent angular velocity of the target, and x be the radial velocity of the scattering point. p Let x be the x-coordinate of the scattering point; then the target's azimuth dimension is expressed as: Then, the ISAR image is normalized as follows: Set threshold T h , will s″ kr (f τ Discretized as a vector s″ kr Then the azimuth sparsity K of the target k Let's call it vector s″ kr Medium greater than T h The number of elements; Let the reference dimension be S. x_ref The reference azimuth resolution is ρ ref Based on the target's azimuth dimensions, the required azimuth resolution for the target is:

3. The method for scheduling networked radar multi-target three-dimensional imaging resources based on deep reinforcement learning as described in claim 2, wherein the threshold T in the first step... s The threshold T is 0.

5. h The value is 0.

2.

4. The resource scheduling method for multi-target three-dimensional imaging of networked radar based on deep reinforcement learning as described in claim 2, wherein the second step specifically comprises: The radar line-of-sight direction vector is determined based on the geographical location of the radar and the target. in, P k P r These are the geographical locations of the target and the radar, respectively; The target's velocity v k The target's radial velocity v is obtained by decomposing the data along the radar line of sight and its orthogonal directions. p With tangential velocity v t The tangential velocity is expressed as: For the target's three-dimensional space, establish a three-dimensional Cartesian coordinate system (O, X, Y, Z) with the target's center O as the origin; for the imaging plane, establish the W-axis along the radar line of sight and the U-axis along the target's tangential velocity direction, with the unit vectors of the two axes being respectively... u u =v t / |v t |, then the radar imaging plane is represented as (O,U,W); Next, the projection relationship is derived. Under far-field conditions, the target ISAR image is considered to be the projection of the target's three-dimensional structure onto the imaging plane; point P is a point outside the imaging plane, and point Q is the projection of point P onto the imaging plane. Two unit vectors perpendicular to a plane are represented as: Let the three-dimensional coordinates of point P be (x, y, z). T The two-dimensional coordinates of point Q are (u, w). T ,but Represented as Equation (6) can be expressed in the form of matrix multiplication: The projection relationship between the three-dimensional coordinates of the scattering point and the two-dimensional coordinates on the imaging plane is then constructed: in, Let be the projection matrix from the three-dimensional coordinates of a point in space to the imaging plane.

5. The method for scheduling networked radar multi-target three-dimensional imaging resources based on deep reinforcement learning as described in claim 4, wherein the third step specifically comprises: Design a size N x ·N y ·N z The three-dimensional Boolean matrix X is used as the optimization variable, where, A matrix value of 1 indicates the presence of a target scattering point at that point. The strong scattering point in the 2D ISAR image of the target formed by the i-th radar is extracted based on the peak intensity of the ISAR image and denoted as G. i The projected image of the target on the imaging plane is represented as F. i (X), where the projection function F i (·) Obtained from the projection matrix φ of the imaging plane; construct the optimization function with the objective of minimizing the difference between the target 2D ISAR image and the projected image: Where, N R This indicates that there are a total of N. R The radar images the target, where λ is the regularization parameter; by solving the optimization function, the three-dimensional scattering point structure of the target can be reconstructed. Then, the impact of radar node distribution on target imaging quality is analyzed. A single two-dimensional image cannot reveal the three-dimensional structure of a target. To locate the three-dimensional scattering points of the target, each target must be imaged by at least two radars. Given two points P1 and P2 in space, and two radars image the target, the projection points of P1 and P2 onto the two imaging planes are Q1 and Q2, respectively. i1 Q i2 Q j1 Q j2 ,set up The normal vector of the imaging plane i is n i θ i for With n i The angle between Q and Q; i1 With Q i2 The distance between them is expressed as: When l i ≥ρ i At that time, P1 and P2 can be distinguished on the imaging plane i, ρ i For the two-dimensional resolution of imaging plane i, the equivalent condition is: If two points can be distinguished at least on an imaging plane, then the correct reconstruction of the two points can be achieved; Therefore, the conditions for achieving a correct reconstruction of the target are expressed as follows: C0 is the set of target scattering points, and N is the number of radars observing the target; this condition is transformed to depend only on the distribution of radar nodes, as follows: Where θ ij ρ is the angle between the two radar imaging planes. rec For the target 3D resolution, θ ij Obtained by the angle between the normal vectors of the two imaging planes; The cross product of the unit vector basis of the imaging plane yields the normal vector n of the i-th and j-th planes. i n j The angle between the normal vectors of the imaging plane is the angle between the imaging planes: If two observation radars meet the above conditions, the target scattering point pair can be distinguished on the imaging plane of at least one of the radars, thus enabling correct reconstruction of the target. Further analysis is conducted on the impact of the number of observation radars on imaging quality.

6. The resource scheduling method for multi-target three-dimensional imaging of networked radar based on deep reinforcement learning as described in claim 5, wherein the third step of further analyzing the impact of the number of observation radars on imaging quality specifically comprises: Let the correct reconstruction result and the incorrect reconstruction result be X and X', respectively. Substitute the reconstruction result into the loss function in the optimization model of equation (9). In order to ensure that the final solution of the model is the correct reconstruction result, the following conditions need to be met: Suppose that X contains two scattering points P1 and P2, and X' contains only one scattering point. How many planes can correctly distinguish between scattering points P1 and P2? P If there are one, then: The left side of equation (15) is all negative, indicating that N P The larger the value, the greater the probability that equation (14) holds true; for most scattering point pairs, the more radars imaging the target, the more imaging planes N that can correctly distinguish them. P The more radars that observe a single target, the lower the probability of the target being incorrectly reconstructed, and the higher the target imaging quality.

7. The method for scheduling networked radar multi-target three-dimensional imaging resources based on deep reinforcement learning as described in claim 5, wherein the fourth step specifically comprises: A multi-target 3D imaging resource scheduling model for networked radars is constructed to allocate radars to each target for 3D imaging, minimizing the consumption of networked radar resources while ensuring a certain level of imaging quality. Taking the shortest total scheduling time as the objective function, and setting constraints for target imaging requirements, we have M radars and N targets to be imaged. An M*N dimensional Boolean matrix W is defined as the optimization variable to represent the scheduling scheme, where W(m,n) = 1 represents the m-th radar imaging n targets. Let T be the total time used by the networked radars to complete multi-target 3D imaging, which is solved through the following steps: Step 4.1 Extract from scheme W which targets each radar needs to observe and record them in the radar target set TSet(m); Step 4.2 Prioritize targets based on their flight speed, distance from radar, and yaw angle. k A comprehensive evaluation is conducted, and then the coherent accumulation time T required for imaging is calculated based on the azimuth resolution required for target imaging. k and observation dimension M k ; Step 4.3 Each radar assigns sparse pulses to the targets in sequence according to priority, until imaging of all targets in the target set is completed; Step 4.4 Calculate the time T taken by each radar to complete two-dimensional ISAR imaging of all targets in its target set. m Then, the longest time is selected as the total time T used by the networked radar to complete multi-target three-dimensional imaging; If steps 4.1-4.4 are defined as Tconsum(W), then the objective function is expressed as: min Tconsum(W) (16) Then determine whether scheme W satisfies the target reconstruction conditions; Assuming the radar observes all targets, calculate the angle between the normal vectors of the imaging planes of every two radars observing N targets, and calculate the minimum angle between the imaging planes of these two radars that must satisfy the reconstruction condition according to equation (12). Record these angles in... A 2D matrix θ RT and θ pho middle, The total number of representatives is Various radar node combination methods; Then, for each column of W, elements are processed according to θ. RT and θ ρ Perform a pairwise bitwise AND operation on each element in sequence to obtain... A matrix RT of dimension 1, where RT(k,n) = 1 indicates that target n is observed by two radars in the k-th radar combination; Then determine whether scheme W satisfies the reconstruction conditions of each objective: Enough=[(RT⊙θ RT )>θ pho ] (17) If any column in Satis contains radar combinations that satisfy the conditions, i.e. It is assumed that scheme W can achieve correct reconstruction of each target; at the same time, in order to reduce the probability of incorrect reconstruction, each target is observed by at least 3 radars. List the resource scheduling models:

8. The resource scheduling method for multi-target three-dimensional imaging of networked radar based on deep reinforcement learning as described in claim 7, wherein the constraints of setting the target imaging requirements in the fourth step are as follows:

1. Each target has at least two radars that can meet the conditions for reconstructing any scattering point, that is, there are two radars that meet the conditions of equation (12); 2. The number of radars that observe the target is as large as possible to reduce the probability of incorrect reconstruction, and each target has at least three radars that image it; 3. Each radar can use a sparse ISAR imaging algorithm based on compressed sensing to observe multiple targets simultaneously, so the pulse sequence allocated to each target needs to meet the sparse reconstruction requirements.

9. The resource scheduling method for multi-target three-dimensional imaging of networked radar based on deep reinforcement learning as described in claim 1, wherein the sixth step specifically comprises: The specific steps for solving the resource scheduling model using the A2C method are as follows: Once the algorithm converges, the optimal radar node scheduling scheme can be obtained. Then, each radar allocates pulses to targets in its target set according to priority based on the scheduling scheme for imaging. Finally, the three-dimensional image of each target is reconstructed based on the three-dimensional imaging algorithm.

Citation Information

Patent Citations

  • Adaptive scheduling method for inverse synthetic aperture radar imaging resources in networking

    CN108761455A

  • Distributed networking radar node position optimization method based on reinforcement learning

    CN116362122A