Distributed MIMO radar multi-target detection method and system
By leveraging the waveform diversity and spatial diversity capabilities of a distributed MIMO radar system, combined with Markov decision processes based on reinforcement learning, and optimizing the beam pattern, the performance degradation of multi-target detection in centralized MIMO radar systems under the target RCS scintillation problem is resolved, achieving more efficient target detection.
Patent Information
- Application Number
- CN202510424702.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-08-01
AI Technical Summary
Centralized MIMO radar systems suffer from performance degradation due to target RCS scintillation, making it difficult to effectively improve multi-target detection performance.
A distributed MIMO radar system is adopted, which utilizes waveform diversity and spatial diversity capabilities, a grid-based GLRT detector and a false target elimination algorithm, and a Markov decision process based on reinforcement learning to optimize the beam pattern to achieve multi-target detection.
It improves the performance and robustness of multi-target detection, prevents targets from being missed, and increases the detection probability and accuracy.
Smart Images

Figure CN120405597A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of MIMO radar target detection, and relates to a multi-target detection method and system for a distributed MIMO radar. Background Art
[0002] As a new type of radar system, the multiple-input multiple-output (MIMO) radar has advantages such as detection ability and strong anti-interference ability compared with the single-input single-output radar system, which has attracted wide attention. MIMO radars can be divided into two categories according to the spacing between the transmitting and receiving antennas: the centralized MIMO radar system and the distributed MIMO radar system; among them, the transmitting and receiving nodes of the centralized MIMO radar system are located at one place and the antennas are closely arranged, with good waveform diversity ability, and the transmitting and receiving nodes of the distributed radar are far apart, with good spatial diversity ability. Although the MIMO radar system can be competent for a variety of complex tasks, target detection has always been one of the most basic functions of the radar system. In order to cope with the multi-target detection task under unknown environmental conditions, due to the advantages that reinforcement learning can solve the model-free problem and does not depend on the prior information of the environment, reinforcement learning has been applied to the centralized MIMO radar system in recent years.
[0003] In the context of the centralized MIMO radar system, its waveform diversity ability is the core on which the multi-target detection method based on reinforcement learning depends when modeling the Markov Decision Process (MDP). Specifically, in the MDP of this type of method, the radar will determine the possible positions of the targets according to the existing detection results, and accordingly optimize the radar transmission waveform, so as to improve the detection performance by concentrating the radar transmission energy in the target direction. However, most of the current research is based on the centralized MIMO radar system, which cannot solve the problem of RCS (radar cross section) scintillation of the target. When the viewing angle of the radar observing the target changes, the target RCS may fade, resulting in a significant reduction in its signal-to-noise ratio. At this time, due to the existence of only a single observation angle for the target in the centralized MIMO radar system, its detection performance often drops severely. The distributed MIMO radar system with spatial diversity ability can effectively solve the above problems. Therefore, how to use the distributed MIMO radar system to improve the multi-target detection performance has become a technical problem to be solved at present. Summary of the Invention
[0004] Aiming at the problems existing in the above traditional technologies, the present invention proposes a multi-target detection method for a distributed MIMO radar and a multi-target detection system for a distributed MIMO radar, which can use the distributed MIMO radar system to improve the multi-target detection performance.
[0005] In order to achieve the above object, the embodiments of the present invention adopt the following technical solutions:
[0006] On the one hand, a multi-target detection method for a distributed MIMO radar is provided, including the steps of:
[0007] Transmit, through the transmitting nodes of the distributed MIMO radar system, a radar transmission signal that has been pre-modulated according to the waveform modulation vector of the transmitting nodes;
[0008] Obtain the radar echo signals received by the receiving nodes of the distributed MIMO radar system;
[0009] Use a grid-based GLRT detector that maximizes the detection probability to detect the radar echo signals, and obtain the detection results of all channels; wherein, the monitoring area of the radar is evenly divided into non-overlapping grids of the same size;
[0010] Use a false target rejection algorithm based on geometric decision-making to reject false targets from the detection results of all channels, obtain the detection results after false target rejection, and calculate the state of reinforcement learning at the current moment; false targets include artifacts and ghosts;
[0011] Select the action of reinforcement learning at the current moment based on the quasi-ε-greedy policy according to the state;
[0012] Calculate the angle that the transmitting nodes need to focus on according to the action, then establish a beam optimization scheme based on the maximum-minimum power gain criterion, and use the semi-definite relaxation method to solve the beam optimization scheme to obtain the optimized waveform modulation vector;
[0013] Calculate the illuminated image of the distributed MIMO radar system according to the beam pattern corresponding to the optimized waveform modulation vector;
[0014] Set the reward of reinforcement learning at the current moment as the sum of the maximum estimated detection probabilities of the detected targets in the illuminated image, update the Q-table of reinforcement learning, and then enter the target detection at the next moment until the detection ends when the maximum number of time steps of reinforcement learning is reached.
[0015] On the other hand, a multi-target detection system for a distributed MIMO radar is also provided, including:
[0016] A signal transmission module, which is used to transmit, through the transmitting nodes of the distributed MIMO radar system, a radar transmission signal that has been pre-modulated according to the waveform modulation vector of the transmitting nodes;
[0017] An echo acquisition module, which is used to obtain the radar echo signals received by the receiving nodes of the distributed MIMO radar system;
[0018] The echo detection module is used to detect the radar echo signal by using a grid-based GLRT detector that maximizes the detection probability, and obtain the detection results of all channels; among them, the monitoring area of the radar is evenly divided into non-overlapping grids of the same size;
[0019] The state calculation module is used to eliminate false targets from the detection results of all channels by using a false target elimination algorithm based on geometric decision-making, obtain the detection results after false target elimination, and calculate the state of reinforcement learning at the current moment; false targets include artifacts and ghosts;
[0020] The action selection module is used to select the action of reinforcement learning at the current moment based on the state according to the quasi-ε-greedy strategy;
[0021] The beam optimization module is used to calculate the angle that the transmitting node needs to focus on according to the action, then establish a beam optimization scheme based on the maximum-minimum power gain criterion, and use the semi-definite relaxation method to solve the beam optimization scheme to obtain the optimized waveform modulation vector;
[0022] The illumination calculation module is used to calculate the illumination image of the distributed MIMO radar system according to the beam pattern corresponding to the optimized waveform modulation vector;
[0023] The reward update module is used to set the reward of reinforcement learning at the current moment as the sum of the maximum estimated detection probabilities of the detected targets in the illumination image, update the Q-table of reinforcement learning, and then enter the target detection at the next moment until the detection ends when the maximum time step of reinforcement learning is reached.
[0024] One of the above technical solutions has the following advantages and beneficial effects:
[0025] In the above distributed MIMO radar multi-target detection method and system, after using a distributed MIMO radar system with both waveform diversity ability and spatial diversity ability and deriving the radar echo signal model, the monitoring area of the radar is divided into multiple grids of the same specification. Based on the maximum detection probability criterion, a grid-based GLRT detector is proposed to detect the radar echo signal for targets. Then, according to the generation mechanism of artifact and ghost targets, a false target elimination method based on geometric decision-making is used to eliminate false targets to achieve a more effective measurement of the number of targets. Next, taking the beam pattern as the optimization object and combining the application requirements of reinforcement learning, a Markov decision process is constructed in the context of a distributed MIMO radar system. According to the maximum-minimum power gain criterion, a general beam optimization scheme is designed, and the corresponding semi-definite relaxation method is used to solve it to obtain the optimized waveform modulation vector, realizing the autonomous focusing of the transmitting energy of the radar on possible targets.
[0026] Compared with the traditional technology, an improved quasi-ε-greedy strategy is adopted in the strategy modeling process to prevent missing targets in view of the detection robustness of the distributed MIMO radar system. Moreover, in the reward modeling process, according to the beam pattern of each transmitting node, the real-time illumination area map of the radar is drawn according to the criterion of the half-beam width. Accordingly, the reward is modeled as the sum of the estimated detection probabilities of the detected targets located in the illuminated area to measure the effectiveness of the beam pattern optimized by the radar, so as to achieve the effect of improving the multi-target detection performance by using the distributed MIMO radar system. Description of the Drawings
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the traditional technology, the following will briefly introduce the drawings required for the description of the embodiments or the traditional technology. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1 It is a flowchart of the multi-target detection method for a distributed MIMO radar in an embodiment;
[0029] Figure 2 It is a schematic structural diagram of the distributed MIMO radar system used in the present invention;
[0030] Figure 3 It is a schematic flowchart of the multi-target detection method for a distributed MIMO radar using reinforcement learning in an embodiment;
[0031] Figure 4 It is a schematic diagram for removing false targets in an embodiment, where Figure 4 (a) is the detection result of the grid GLRT detector, Figure 4 (b) is from Figure 4 (a) The detection result after removing the artifact targets, Figure 4 (c) is from Figure 4 (b) The detection result after removing the ghost targets;
[0032] Figure 5 It is a schematic diagram of the simulation scenario setting in an embodiment;
[0033] Figure 6 It is a line graph of the detection probabilities of the proposed method and the adaptive method in two static scenarios in an embodiment, where Figure 6 (a) is the comparison result of static scenario 1, Figure 6 (b) is the comparison result of static scenario 2;
[0034] Figure 7It is a block diagram of a distributed MIMO radar multi-target detection system in an embodiment. Detailed implementation manners
[0035] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.
[0036] It should be noted that referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Displaying this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art can understand that the embodiments described herein can be combined with other embodiments. The term "and / or" used in the description and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0037] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings in the embodiments of the present invention.
[0038] Reinforcement learning is a machine learning method. It can achieve autonomous perception of the environment by repeatedly executing a pre-modeled Markov Decision Process (MDP), and use the gradually accumulated experience to help the agent take correct decisions to complete the preset tasks. MDP mainly includes four elements: state, policy, action and reward. According to whether the modeled state and action are discrete, the reinforcement learning algorithm can be divided into a tabular solution method and a tabular approximate solution method.
[0039] A distributed MIMO radar system with spatial diversity ability can effectively solve the RCS flickering problem of targets existing in a centralized MIMO radar system. The distributed MIMO radar system has multiple transmitting nodes and receiving nodes with relatively large distances from each other. Each transmitting node and receiving node form a channel, thereby enabling simultaneous multi-view observation of each target. Therefore, the present invention intends to introduce a distributed MIMO radar system to implement distributed MIMO radar multi-target detection based on reinforcement learning, so as to improve the multi-target detection performance.
[0040] In one embodiment, as Figure 1As shown, a method for multi-target detection of a distributed MIMO radar is provided, which may include the following steps S10 to S24:
[0041] S10, transmitting, through a transmitting node of a distributed MIMO radar system, a radar transmission signal pre-modulated according to a waveform modulation vector of the transmitting node;
[0042] S12, acquiring a radar echo signal received by a receiving node of the distributed MIMO radar system;
[0043] S14, using a grid-based GLRT detector that maximizes the detection probability to detect the radar echo signal, and obtaining detection results of all channels; wherein, the monitoring area of the radar is evenly divided into non-overlapping grids of the same size;
[0044] S16, using a false target rejection algorithm based on geometric decision-making to reject false targets from the detection results of all channels, obtaining the detection results after false target rejection, and calculating the state of reinforcement learning at the current moment; false targets include artifacts and ghosts;
[0045] S18, selecting an action of reinforcement learning at the current moment based on the quasi-ε-greedy policy according to the state;
[0046] S20, calculating the angle that the transmitting node needs to focus on according to the action, then establishing a beam optimization scheme based on the maximum-minimum power gain criterion, and using the semi-definite relaxation method to solve the beam optimization scheme to obtain an optimized waveform modulation vector;
[0047] S22, calculating an illumination image of the distributed MIMO radar system according to the beam pattern corresponding to the optimized waveform modulation vector;
[0048] S24, setting the reward of reinforcement learning at the current moment as the sum of the maximum estimated detection probabilities of the detected targets in the illumination image, updating the Q-table of reinforcement learning, and then entering the target detection at the next moment until the detection ends when the maximum number of time steps of reinforcement learning is reached.
[0049] The above-mentioned distributed MIMO radar multi-target detection method, after using a distributed MIMO radar system with both waveform diversity and spatial diversity capabilities and deriving the radar echo signal model, divides the monitoring area of the radar into multiple grids of equal specifications. Based on the maximum detection probability criterion, a grid-based GLRT detector is proposed to detect the radar echo signal for targets. Then, according to the generation mechanism of artifacts and ghost targets, a false target elimination method based on geometric decision-making is used to eliminate false targets to achieve a more effective measurement of the number of targets. Next, taking the beam pattern as the optimization object and combining the application requirements of reinforcement learning, a Markov decision process is constructed in the context of a distributed MIMO radar system. According to the maximum minimum power gain criterion, a general beam optimization scheme is designed, and the corresponding semi-definite relaxation method is used to solve for the optimized waveform modulation vector to achieve the autonomous focusing of the radar's transmitted energy on possible targets.
[0050] Compared with traditional technologies, an improved quasi-ε-greedy strategy is adopted in the policy modeling process in view of the detection robustness of the distributed MIMO radar system to prevent missing targets. Moreover, in the reward modeling process, according to the beam pattern of each transmitting node and in accordance with the criterion of the half-beam width, a real-time illumination area map of the radar is drawn. Based on this, the reward is modeled as the sum of the estimated detection probabilities of the detected targets located within the illuminated area to measure the effectiveness of the beam pattern optimized by the radar, achieving the effect of improving the multi-target detection performance using the distributed MIMO radar system.
[0051] It can be understood that the distributed MIMO radar system used can be as Figure 2 shown. There are a total of M transmitting nodes and N receiving nodes in the radar system. Each node is equipped with Q closely arranged antennas, and the spacing between the antennas is half a wavelength. Denote the base signal connected to the m-th (m = 1, 2, …, M) transmitting node as Φ m (t), t ∈ [0, T p , where T p is the pulse duration. The base signals {Φ m (t)|m = 1, 2, …, M} are orthogonal to each other. For the m-th transmitting node, a separate phase modulation coefficient is set for each of its antennas. Then, a waveform modulation vector of length Q can be formed by the phase modulation coefficients of the Q antennas, which can be denoted as Then the transmitted signal of the m-th transmitting node is expressed as:
[0052]
[0053] And the transmitted energy of the m-th transmitting node
[0054] As Figure 3As shown in the figure, first, transmit the pre-modulated radar transmit signal and receive the radar echo signal. Assume that there are a total of C targets in the monitoring area of the radar. Without considering time delay for the moment, the radar signal received at the c-th (c ∈ [1, C]) target transmitted by the m-th (m ∈ [1, M]) transmit node is expressed as:
[0055]
[0056] where α mc is a complex coefficient representing the propagation loss of the radar signal in the transmission path (from the m-th transmit node to the c-th target), θ mc represents the angle of the c-th target relative to the m-th transmit node, and a T (·) represents the transmit steering vector, and there is:
[0057]
[0058] The superscript in the square brackets in Equation (3) represents the transpose. Then, the radar echo signal received at the n-th receive node can be expressed as:
[0059]
[0060] where α c , α nc are both complex coefficients, representing the radar cross-section area of the c-th target and the propagation loss of the radar echo signal in the receiving path (from the c-th target to the n-th receive node) respectively. Denote α mnc =α c α nc α mc . represents the angle of the c-th target relative to the n-th receive node, and the definition of a R (·) is the same as that of a T (·). τ mnc represents the time delay. represents the noise signal.
[0061] Subsequently, the signals contained in y n (t) (i.e., s m (t), m = 1, 2, …, M) can be orthogonally separated by using matched filtering, and then the signal is sampled according to the sampling frequency f s . Assume represents rounding up to the nearest integer, that is, there are K range cells.
[0062] Assume that there are a total of L channels in the distributed MIMO radar system, and If the c-th target is located at the On the k-th range cell of the n-th channel, the corresponding sampled vector can be expressed as where l m and l n represent the numbers of the transmitting node and the receiving node corresponding to the l-th channel respectively; otherwise, is the signal obtained by matched filtering of w.
[0063] Then, a grid-based GLRT (Generalized Likelihood Ratio Test) detector is used to detect the received radar echo signal.
[0064] The binary hypothesis description of the radar echo detection problem is as follows:
[0065]
[0066] where, a R and a T are abbreviations of a T (θ) respectively.
[0067] Without loss of generality, w follows a circularly symmetric complex Gaussian distribution, and the variance of w is denoted as After calculation, the GLRT statistic Λ corresponding to the above formula is expressed as:
[0068]
[0069] The detection threshold is represents the inverse of the cumulative distribution function of, and the estimated detection probability is expressed as:
[0070]
[0071] To determine the position of the target, the monitoring area of the radar is evenly divided into non-overlapping grids of the same size. Assume that the monitoring area of the radar is a square area in the two-dimensional plane (denoted as ), then each point in the square area can be uniquely represented by the coordinates {X s , Y s}. Assume that the square area can be divided into G non-overlapping grids of size ω x ×ω y , and the grid center coordinates of each grid are denoted as Then the two-dimensional plane area covered by the g-th (g ∈ [1, G]) grid can be mathematically expressed as:
[0072]
[0073] Suppose the set of range cells occupied by the \(g\)-th grid on the \(l\)-th channel is: <·> represents the number of samples in the set. There are GLRT statistics for the \(g\)-th grid on the \(l\)-th channel, and these GLRT statistics are in the form of:
[0074]
[0075] where \(l\) n points to the receiving node of the \(l\)-th channel, and its coordinate position is denoted as represents the receiving steering vector of the grid center relative to the receiving node of the \(l\)-th channel, expressed as:
[0076]
[0077] In addition, represents the radar echo data at the -th range cell of the \(l\)-th channel.
[0078] To maximize the detection probability and achieve the purpose of not missing targets, the grid-based GLRT detector is constructed in the following form:
[0079]
[0080] where, represents that there is no target in the current grid, and the alternative hypothesis represents that there is a target in the current grid. Further, the following formula (12) is called the estimated detection probability of the \(g\)-th grid, where \(Q1()\) represents the first-order Marcum function.
[0081]
[0082] Next, obtain the state at time \(j\), denoted as \(s\) j .
[0083] Denote the detection result of the \(l\)-th channel through the grid GLRT detector as Then the detection results of all channels (in vector form) are represented as:
[0084]
[0085] where
[0086] To achieve multi-target detection tasks, it is necessary to determine the number of targets according to ρ. However, due to the system characteristics of distributed MIMO radars, various false targets will be included in ρ. Therefore, the focus of the next step of processing is to eliminate false targets.
[0087] In the detection scenario of a distributed MIMO radar system, false targets can be divided into three categories: artifacts, ghosts, and false alarm targets. Artifacts are caused by the energy echo reflection of a target on the same range cell, ghosts are caused by the intersection of range cells of different targets on different channels, and false alarm targets are generated when the GLRT statistical value of noise exceeds the threshold. Obviously, the appearance positions of artifacts and ghosts depend on real targets, while false alarm targets appear randomly. The focus of this specification is to eliminate artifacts and ghost targets.
[0088] From the formation mechanism of artifacts, it can be seen that they appear in the same range cell as the target. Since in a distributed MIMO radar system, the distribution positions of range cells in different channels are different, real targets should exist at the intersection of artifacts in different channels. Accordingly, artifacts can be further eliminated using the following formula:
[0089]
[0090] Denote as the detection result (in vector form) after eliminating artifacts from ρ, where 1 represents the existence of a target in the current grid and 0 represents noise or clutter in the current grid.
[0091] Subsequently, the fast connected-component labeling algorithm (FCCL) is used to preliminarily determine the number of targets. This step can be expressed as:
[0092]
[0093] where represents the function operation of FCCL, is the detection result after eliminating artifacts (g = 1, 2,..., G) forms a binary image by assigning 0 or 1 values according to its grid position in the monitoring area. W is the number of connected regions in , denote as the th connected region in
[0094] In one embodiment, regarding step S16 described above, during the process of removing false targets from the detection results of all channels using the false target removal algorithm based on geometric decision-making, the ghost decision condition is used to determine whether the connected region in the binary image formed by the detection results after removing artifacts is a ghost target; the ghost decision condition is as follows:
[0095] If there exists a connected region that is a subset of the set of range cells occupied by two other different connected regions on at least 2 channels respectively, then the connected region is a ghost target; otherwise, it is not a ghost target.
[0096] Specifically, from the formation mechanism of ghosts, Equation (14) is ineffective in removing ghosts. To achieve the goal of removing ghosts, further processing is required. First, denote the grid set included in the connected region as denotes the number of grids included in the range cells occupied by
[0097]
[0098] in the l-th channel. From the formation mechanism of ghosts, the ghost decision condition can be determined: If there exists a connected region that is a subset of the set of range cells occupied by two other different connected regions on at least 2 channels respectively, then the connected region is a ghost target. This ghost decision condition can be described in mathematical language as:
[0099]
[0100] then belongs to ghosts. Assume that it is known that is the set of connected regions where ghosts are located, and Z is the number of ghost regions. Then the grid set included in the z-th ghost region is Furthermore, it can be set that:
[0101]
[0102] then is the detection result obtained by removing ghosts from , where 1 represents the presence of a target in the current grid, 0 represents the current grid being noise or clutter. Denote as the binary label image constructed from .
[0103] All in all, the false target removal algorithm based on geometric decision-making from the detection results of each radar channel can be as shown in Algorithm 1 in Table 1. The schematic diagram of removing artifacts and ghost targets can be as shown in Figure 4 shown, whereFigure 4 (a) is the detection result of the grid GLRT detector, Figure 4 (b) is from Figure 4 (a) after removing the artifact targets, Figure 4 (c) is from Figure 4 (b) after removing the ghost targets. The output item of Algorithm 1 is the finally obtained target area, and the defined state s j is:
[0104]
[0105] Table 1
[0106]
[0107] Furthermore, according to the policy, obtain the action at time j, denoted as a j .
[0108] Denote the maximum number of targets that the radar system can monitor simultaneously as Mo, and define the action space as: At time j, first construct a temporary action space according to the state s j :
[0109]
[0110] where,
[0111]
[0112] is a binary label image constructed by ρ. The action a j is selected by the quasi-ε-greedy policy of the following formula:
[0113]
[0114] where, means randomly select an action from the set.
[0115] Then, optimize the beam modulation vectors of each transmitting node.
[0116] First, determine the positions of a j targets. The positions of I targets can already be determined from . Denote the grid occupied by the center position of each grid as Regard the center position of each connected region as the target position, denoted as i = 1, 2, …, I, then the positions of the first I targets can be obtained by the following formula:
[0117]
[0118] Then, search for the remaining a from among the j -s j targets. Specifically, denote the grid occupied by Ω v (v = 1, 2, …, V) as The central position of each grid is (u = 1, 2, …, >Ω v )). Denote the central position of Ω v as Then:
[0119]
[0120] Assume that there is at most one real target in a connected region. Then, if Ω v (v = 1, 2, …, V) satisfies:
[0121]
[0122] then it is not considered that there is a hidden target in Ω v .
[0123] In one embodiment, regarding the above step S20, further, in the process of calculating the angle that the transmitting node needs to focus on according to the action, the estimated detection probability is used to select the connected region where a hidden target is most likely to exist.
[0124] Specifically, the estimated detection probability is used to select the connected region where a hidden target is most likely to exist. Denote the maximum estimated detection probability (denoted as v ) of the grids in Ω as:
[0125]
[0126] Based on this, is sorted to find the highest a j -s j connected regions corresponding to the detection probabilities, which are regarded as having hidden targets.
[0127] Uniformly denote the central position of the a-th target as Then its angle relative to the m-th transmitting node is expressed as:
[0128]
[0129] The following gives an optimization scheme for the beam modulation vector and its solution method.
[0130] Specifically, taking the m-th transmitting node as an example, let its waveform modulation vector Transmission energy Angle Action The angles that need to be focused are θ1, θ2, …, θ A , then the optimization scheme of the beam modulation vector can be described as:
[0131]
[0132] Introduce auxiliary variables And denote Then the above optimization scheme can be transformed into a beam optimization scheme based on the maximum-minimum power gain criterion:
[0133]
[0134] Among them, constraint condition (i) is because H is positive semi-definite, and constraint condition (iii) comes from the energy limitation of the transmitting node.
[0135] Next, use the semi-definite relaxation method to solve the optimization scheme shown in Equation (29). First, without considering constraint condition (ii), then Equation (29) degenerates into a convex optimization problem:
[0136]
[0137] It can be directly solved using the CVX toolbox (a MATLAB toolbox for convex optimization). Denote the optimized solution as For Perform eigenvalue decomposition, expressed as:
[0138]
[0139] Among them, U = [u1, u2, …, u Q is a unitary matrix, u q (q ∈ [1, Q]) represents the q-th column of matrix U, Λ is a diagonal matrix composed of the absolute values of the eigenvalues of , and the superscript H in the upper right corner of matrix U represents conjugation. Generally, let the eigenvalues in Λ be arranged in descending order.
[0140] If Then the solution of Equation (29) is:
[0141] If Use the Gaussian random optimization method (an existing method) to obtain a sub-optimal solution of Equation (29). This process can be described as randomly taking E vectors with a fixed length from the complex unit circle, calculating their objective function values respectively. Subsequently, find the random vector corresponding to the maximum value and calculate the sub-optimal solution.
[0142] The specific optimization algorithm is shown in Algorithm 2 of Table 2, where lines 8 - 12 are the process of Gaussian random optimization, and line 12 is to meet the constraint of transmission energy.
[0143] Table 2
[0144]
[0145] Next, calculate the reward at time j, denoted as r j 。
[0146] First, borrow the idea of the half - beamwidth to define the angular range of the beam pattern focusing. Specifically, assume that a certain peak (the power gain is denoted as B0) of the beam pattern of the l - th channel is known to be located at the angle θ0. Then, the angular interval centered at θ0 and including the angles where the power gain is greater than B0 / 2 is called the focusing range of the beam pattern at θ0. Further, if the center position of the g - th grid satisfies:
[0147]
[0148] then the grid is said to be illuminated by the transmitting node of the l - th channel, denoted as otherwise denoted as
[0149] In practice, the optimized beam pattern may have the following two problems: (1) The peak power gain of the beam pattern does not necessarily lie in θ1, θ2, …, θ I ; (2) There may be several adjacent focusing angles co - existing in the same focusing range.
[0150] The following describes the method (using the existing method) to solve the above problems taking the l - th channel as an example: First step, first find all the maxima of the beam pattern of the m - th transmitting node, and arrange all the maxima in a descending order into a vector: D is the number of maxima, and the angles at which all the maxima are located are respectively Then the peak of the beam pattern must be included in the vector However, the vector will also include sidelobe peaks. Second step, use the gradient - descent method to obtain a dynamic threshold and adaptively find the peak at which the radar in the l - th channel vector truly focuses. Construct the gradient vector as shown in Equation (33) below:
[0151]
[0152] The larger the element in Equation (33) indicates the more significant the difference between two adjacent values in
[0153]
[0154] Then the adaptive threshold is taken as The actual focusing angle of the radar is The corresponding power gain is
[0155] Based on this, the binary label of the g-th grid in the monitoring area of the m-th transmitting node is calculated according to the following formula (35): Assignment:
[0156]
[0157] By The binary label image is called the illuminated image of the emission node of the lth channel in the monitoring area.
[0158] Furthermore, let η g Indicates whether the g-th grid is illuminated by the beam pattern of a certain channel, then:
[0159]
[0160] Let η=[η1,η2,…,η G ] T The binary label image constructed by η is called the illumination image of the radar system on the monitoring area.
[0161] The reward in this specification is designed to be the sum of the maximum estimated detection probabilities of the detected objects in the illuminated image, as follows:
[0162]
[0163] The above formula (37) ignores the description of the moment. At time j, η (·) Located at time j-1.
[0164] Summarizing equations (32)-(36), an algorithm for calculating the illuminated image of a radar system is given, as shown in Algorithm 3 of Table 3.
[0165] Table 3
[0166]
[0167] Then, update the size as follows Q form:
[0168] Q(s j-1 ,a j-1 )←Q(s j-1 ,a j-1 )+α·[r j+γ·Q(s j ,a j )-Q(s j-1 ,a j-1 )](38)
[0169] In Equation (38), Q(s, a) represents the value of taking action in state , α is the learning rate, γ is the discount factor, and is the set of states.
[0170] Finally, if j = J (J is the maximum number of time steps), break out of the loop; otherwise, go to the next time step j + 1 and jump to the previous step of transmitting the pre-modulated radar transmission signal.
[0171] Summarizing all the above processing steps, the algorithm flow of the above distributed MIMO radar multi-target detection method can be as shown in Algorithm 4 of Table 4 below, where the Q-table is the core data structure of the Q-learning algorithm in reinforcement learning and is a table for storing Q-values (state-action values).
[0172] Table 4
[0173]
[0174] In some embodiments, as Figure 5 shown, this experimental example considers two station layout schemes for distributed MIMO radar systems, both of which include 2 transmitting nodes and 2 receiving nodes, for a total of 4 channels.
[0175] Station layout scheme 1: The two transmitting nodes are located at {10 km, 0 km} and {90 km, 0 km} respectively, and the two receiving nodes are located at {40 km, 0 km} and {120 km, 0 km} respectively.
[0176] Station layout scheme 2: The two transmitting nodes are located at {10 km, 0 km} and {60 km, 0 km} respectively, and the two receiving nodes are located at {20 km, 0 km} and {120 km, 0 km} respectively.
[0177] The four corner coordinates of the square monitoring area are uniformly {38 km, 52 km}, {38 km, 89 km}, {76 km, 52 km}, and {76 km, 89 km}.
[0178] The radar parameters are set as follows: The radar carrier frequency is 0.8 GHz, the sampling frequency is 1 MHz, and the pulse width is 5 ms. Then the radar has a total of 5000 range cells, and the range resolution is 150 meters. In this example, let ω x = ω y= 150 m, the entire monitoring area can be divided into 253 and 247 grids horizontally and vertically respectively. The energy of each transmitting node of the radar is 1, that is Clutter power
[0179] The parameter settings of reinforcement learning are as shown in Table 5 below, where the sizes of the state and action spaces are both set to 6, that is
[0180] Table 5
[0181]
[0182] Set the following two scenarios.
[0183] Static Scenario 1: Station layout plan 1; There are 3 targets in the monitoring area. Target 1 is located at {45 km, 82 km}, Target 2 is located at {50 km, 62 km}, Target 3 is located at {72 km, 73 km}, and the target signal-to-clutter ratios are all -14 dB; The number of antennas Q of each node is 32.
[0184] Static Scenario 2: Station layout plan 2; There are 3 targets in the monitoring area. Target 1 is located at {45 km, 82 km}, Target 2 is located at {50 km, 62 km}, Target 3 is located at {72 km, 73 km}, and the target signal-to-clutter ratios are all -14 dB; The number of antennas Q of each node is 32.
[0185] After calculation, Targets 1, 2, and 3 are respectively in the grids with indexes {47, 200}, {80, 67}, and {227, 140}.
[0186] Set the false alarm rate to 10 -4 , After 1000 Monte Carlo tests, the Figure 6 shown experimental results are obtained. This figure compares the detection probability broken line graphs of the proposed method (i.e., the above-mentioned distributed MIMO radar multi-target detection method) and the existing adaptive method in the above two static scenarios, that is, the change curve of the detection probability with time. Among them, the adaptive method is the version of the proposed method without reinforcement learning, which can be understood as the proposed method where the action is always equal to the state, that is s j = a j
[0187] Figure 6It is shown that the detection probabilities of the proposed method and the adaptive method both increase rapidly within \(j\in[0,10]\), and basically tend to be stable after \(j > 10\). Among them, the improvement degree of the detection performance of the adaptive method is limited because it only optimizes the beam pattern of each transmitting node according to the detected targets, and it is easy to miss targets in this way. Driven by reinforcement learning, the proposed method can regularly explore the areas of undetected targets on the basis of the detected targets. Therefore, the detection performance of the proposed method improves significantly faster than that of the adaptive method, and the final detection probability is also significantly higher than that of the adaptive detection method.
[0188] It should be understood that although the various steps in the above process Figure 1 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above process Figure 1 may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0189] In one embodiment, as Figure 7As shown in the figure, a distributed MIMO radar multi-target detection system 100 is provided, including a signal transmission module 11, an echo acquisition module 13, an echo detection module 15, a state calculation module 17, an action selection module 19, a beam optimization module 21, an illumination calculation module 23, and a reward update module 25. Among them, the signal transmission module 11 is used to transmit a radar transmission signal modulated in advance according to the waveform modulation vector of the transmission node through the transmission nodes of the distributed MIMO radar system. The echo acquisition module 13 is used to acquire the radar echo signals received by the receiving nodes of the distributed MIMO radar system. The echo detection module 15 is used to detect the radar echo signals by using a grid-based GLRT detector that maximizes the detection probability, and obtain the detection results of all channels; among them, the monitoring area of the radar is evenly divided into non-overlapping grids of the same size. The state calculation module 17 is used to eliminate false targets from the detection results of all channels by using a false target elimination algorithm based on geometric decision-making, obtain the detection results after false target elimination, and calculate the state of reinforcement learning at the current moment; false targets include artifacts and ghosts. The action selection module 19 is used to select the action of reinforcement learning at the current moment based on the state according to the quasi-ε-greedy strategy. The beam optimization module 21 is used to calculate the angles that the transmission nodes need to focus on according to the action, then establish a beam optimization scheme based on the maximum-minimum power gain criterion, and use the semi-definite relaxation method to solve the beam optimization scheme to obtain the optimized waveform modulation vector. The illumination calculation module 23 is used to calculate the illumination image of the distributed MIMO radar system according to the beam pattern corresponding to the optimized waveform modulation vector. The reward update module 25 is used to set the reward of reinforcement learning at the current moment as the sum of the maximum estimated detection probabilities of the detected targets in the illumination image, update the Q-table of reinforcement learning, and then enter the target detection at the next moment until the detection ends when the maximum time step of reinforcement learning is reached.
[0190] For the above-mentioned distributed MIMO radar multi-target detection system 100, after using a distributed MIMO radar system with both waveform diversity and spatial diversity capabilities and deriving the radar echo signal model, the monitoring area of the radar is divided into multiple grids of the same specification. After detecting the radar echo signals by using a grid-based GLRT detector based on the maximum detection probability criterion, false targets are eliminated by using a false target elimination method based on geometric decision-making according to the generation mechanism of artifacts and ghost targets, so as to achieve a more effective measurement of the number of targets. Then, taking the beam pattern as the optimization object and combining the application requirements of reinforcement learning, a Markov decision process is constructed in the context of the distributed MIMO radar system. A general beam optimization scheme is designed according to the maximum-minimum power gain criterion, and the corresponding semi-definite relaxation method is used to solve it to obtain the optimized waveform modulation vector, realizing the autonomous focusing of the radar transmission energy on possible targets.
[0191] Compared with the traditional technology, an improved quasi-ε-greedy strategy is adopted in the strategy modeling process to prevent missing targets in view of the detection robustness of the distributed MIMO radar system. Moreover, in the reward modeling process, according to the beam pattern of each transmitting node, the real-time illuminated area map of the radar is drawn according to the criterion of half beam width. Accordingly, the reward is modeled as the sum of the estimated detection probabilities of the detected targets located in the illuminated area to measure the effectiveness of the optimized beam pattern of the radar, so as to achieve the effect of improving the multi-target detection performance by using the distributed MIMO radar system.
[0192] In one embodiment, in the process of removing false targets from the detection results of all channels by using the false target removal algorithm based on geometric decision-making, the ghost judgment condition is used to determine whether the connected region in the binary image formed by the detection results after removing artifacts is a ghost target; the ghost judgment condition is:
[0193] If there exists a connected region that is a subset of the set of range cells occupied by two other different connected regions on at least 2 channels respectively, then the connected region is a ghost target, otherwise it is not a ghost target.
[0194] In one embodiment, in the process of calculating the angle that the transmitting node needs to focus on according to the action, the estimated detection probability is used to select the connected region most likely to have hidden targets.
[0195] In one embodiment, in the process of removing false targets from the detection results of all channels by using the false target removal algorithm based on geometric decision-making, the following formula is used to remove artifacts:
[0196]
[0197] where, represents the detection result after removing artifacts from the detection results ρ of all channels, g ∈ [1, G] and G represents the number of divided grids, and the superscript T in the upper right corner represents transpose; ρ g represents the detection result of the g-th grid, 1 represents that there is a target in the current grid, and 0 represents that the current grid is noise or clutter.
[0198] In one embodiment, in the process of removing false targets from the detection results of all channels by using the false target removal algorithm based on geometric decision-making, the following formula is used to remove ghosts:
[0199]
[0200] where, represents the detection result obtained by removing ghosts from , represents the set of grids included in the connected region , Denote the set of grids included in the z-th ghost region, W denote the number of connected regions in the binary image formed by the detection results after removing artifacts, Z denote the number of ghost regions, 1 represents that there is a target in the current grid, and 0 represents that the current grid is noise or clutter.
[0201] Regarding the specific limitations of a distributed MIMO radar multi-target detection system 100, reference can be made to the corresponding limitations of a distributed MIMO radar multi-target detection method in the above text, which will not be elaborated here.
[0202] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), memory bus dynamic random access memory (Rambus DRAM, abbreviated as RDRAM), and interface dynamic random access memory (DRDRAM), etc.
[0203] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0204] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it cannot be understood as a limitation on the protection scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, which all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A distributed MIMO radar multi-target detection method, characterized in that, Including the steps: Transmitting, by a transmitting node of a distributed MIMO radar system, a radar transmission signal pre-modulated according to the waveform modulation vector of the transmitting node; Obtaining the radar echo signal received by a receiving node of the distributed MIMO radar system; Detecting the radar echo signal by using a grid-based GLRT detector that maximizes the detection probability to obtain the detection results of all channels; wherein, the monitoring area of the radar is evenly divided into non-overlapping grids of the same size; Eliminating false targets from the detection results of all channels by using a false target elimination algorithm based on geometric decision-making to obtain the detection results after false target elimination and calculating the state of reinforcement learning at the current moment; the false targets include artifacts and ghosts; Selecting the action of reinforcement learning at the current moment based on the quasi-ε-greedy policy according to the state; Calculating the angle that the transmitting node needs to focus on according to the action, then establishing a beam optimization scheme based on the maximum-minimum power gain criterion, and solving the beam optimization scheme by using the semi-definite relaxation method to obtain the optimized waveform modulation vector; Calculating the illuminated image of the distributed MIMO radar system according to the beam pattern corresponding to the optimized waveform modulation vector; Setting the reward of reinforcement learning at the current moment as the sum of the maximum estimated detection probabilities of the detected targets in the illuminated image, updating the Q-table of reinforcement learning, and then entering the target detection at the next moment until the detection ends when the maximum number of time steps of reinforcement learning is reached.
2. The distributed MIMO radar multi-target detection method according to claim 1, wherein, During the process of eliminating false targets from the detection results of all channels by using a false target elimination algorithm based on geometric decision-making, using a ghost judgment condition to determine whether the connected region in the binary image formed by the detection results after eliminating artifacts is a ghost target; the ghost judgment condition is: If there exists a connected region that is a subset of the set of range cells occupied by two other different connected regions on at least 2 channels respectively, then the certain connected region is a ghost target, otherwise it is not a ghost target.
3. The distributed MIMO radar multi-target detection method according to claim 1 or 2, characterized in that During the process of calculating the angle that the transmitting node needs to focus on according to the action, using the estimated detection probability to select the connected region most likely to have hidden targets.
4. The distributed MIMO radar multi-target detection method according to claim 3, characterized in that, During the process of eliminating false targets from the detection results of all channels by using a false target elimination algorithm based on geometric decision-making, eliminating artifacts by using the following formula: Among them, represents the detection result after removing artifacts from the detection results ρ of all channels, where g ∈ [1, G] and G represents the number of divided grids, and the superscript T represents transpose; ρg represents the detection result of the g-th grid, 1 indicates that there is a target in the current grid, and 0 indicates that the current grid is noise or clutter.
5. The distributed MIMO radar multi-target detection method according to claim 3, wherein During the process of eliminating false targets from the detection results of all channels by using a false target elimination algorithm based on geometric decision-making, eliminating ghosts by using the following formula: Among them, represents the detection result obtained by removing ghost images from , represents the set of grids included in the connected region , represents the set of grids included in the z-th ghost region. W represents the number of connected regions in the binary image composed of the detection results after removing artifacts. Z represents the number of ghost regions. 1 represents that there is a target in the current grid, and 0 represents that the current grid is noise or clutter.
6. A distributed MIMO radar multi-target detection system, characterized in that, Including: A signal transmission module, configured to transmit, by a transmitting node of a distributed MIMO radar system, a radar transmission signal pre-modulated according to the waveform modulation vector of the transmitting node; An echo acquisition module, configured to obtain the radar echo signal received by a receiving node of the distributed MIMO radar system; An echo detection module, configured to detect the radar echo signal by using a grid-based GLRT detector that maximizes the detection probability to obtain the detection results of all channels; wherein, the monitoring area of the radar is evenly divided into non-overlapping grids of the same size; A state calculation module, which is used to eliminate false targets from the detection results of all channels by using a false target elimination algorithm based on geometric decision-making, obtain the detection results after false target elimination, and calculate the state of reinforcement learning at the current moment; the false targets include artifacts and ghost images. An action selection module, which is used to select the action of reinforcement learning at the current moment based on the quasi-ε-greedy policy according to the state. A beam optimization module, which is used to calculate the angle that the transmitting node needs to focus on according to the action, then establish a beam optimization scheme based on the maximum-minimum power gain criterion, and use the semi-definite relaxation method to solve the beam optimization scheme to obtain the optimized waveform modulation vector. An illumination calculation module, which is used to calculate the illumination image of the distributed MIMO radar system according to the beam pattern corresponding to the optimized waveform modulation vector. A reward update module, which is used to set the reward of reinforcement learning at the current moment as the sum of the maximum estimated detection probabilities of the detected targets in the illumination image, update the Q-table of reinforcement learning, and then enter the target detection at the next moment until the detection ends when the maximum number of time steps of reinforcement learning is reached.
7. The distributed MIMO radar multi-target detection system according to claim 6, characterized in that During the process of eliminating false targets from the detection results of all channels by using a false target elimination algorithm based on geometric decision-making, the ghost image decision condition is used to determine whether the connected region in the binary image formed by the detection results after artifact elimination is a ghost target; the ghost image decision condition is: If there exists a connected region that satisfies being a subset of the set of range cells occupied by another two different connected regions on at least 2 channels respectively, then the certain connected region is a ghost target, otherwise it is not a ghost target.
8. The distributed MIMO radar multi-target detection system according to claim 6 or 7, characterized in that During the process of calculating the angle that the transmitting node needs to focus on according to the action, the estimated detection probability is used to select the connected region most likely to have hidden targets.
9. The distributed MIMO radar multi-target detection system according to claim 8, characterized in that, During the process of eliminating false targets from the detection results of all channels by using a false target elimination algorithm based on geometric decision-making, the following formula is used to eliminate artifacts: Among them, represents the detection result after removing artifacts from the detection results ρ of all channels. g ∈ [1, G], where G represents the number of divided grids, and the superscript T represents transpose; ρg represents the detection result of the g-th grid, 1 indicates that there is a target in the current grid, and 0 indicates that the current grid is noise or clutter.
10. The distributed MIMO radar multi-target detection system according to claim 8, characterized in that, During the process of eliminating false targets from the detection results of all channels by using a false target elimination algorithm based on geometric decision-making, the following formula is used to eliminate ghost images: Among them, represents the detection result obtained by removing ghost images from , represents the set of grids included in the connected region , represents the set of grids included in the z-th ghost region. W represents the number of connected regions in the binary image composed of the detection results after removing artifacts. Z represents the number of ghost regions. 1 represents the existence of a target in the current grid, and 0 represents that the current grid is noise or clutter.