Multi-radar cooperative detection target distribution method for multiple targets
By adopting layered deep reinforcement learning and predefined action pool methods in the multi-radar collaborative detection system, the problem of multi-radar collaborative detection target allocation in complex dynamic environments is solved, and efficient target tracking and resource utilization are achieved.
Patent Information
- Application Number
- CN202510337126.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
AI Technical Summary
The problem of multi-radar collaborative detection target allocation is difficult to effectively solve in complex dynamic environments, especially in the complexity of high-mobility targets, limited resources and uncertain environments.
The multi-radar multi-flight target collaborative detection target allocation method based on layered deep reinforcement learning is adopted. By combining action pool predefined and dual-layer strategy collaborative optimization, the engineering constraints of the radar system are explicitly integrated to reduce the complexity of high-dimensional discrete action space.
It has achieved the improvement of target tracking efficiency and resource utilization in multi-radar systems, and can effectively allocate tracking targets in complex dynamic environments, improving detection performance, system efficiency and coordination efficiency.
Smart Images

Figure CN120218537A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-radar collaborative detection decision-making, and specifically relates to a multi-radar multi-flying target collaborative detection target allocation method based on hierarchical deep reinforcement learning, which is applicable to target allocation and collaborative optimization of multi-radar systems in complex dynamic environments. Background Art
[0002] With the rapid evolution of science and technology, traditional single-radar detection means are difficult to meet the requirements of real-time, accuracy, and continuity in the face of multi-batch and highly maneuverable targets. To effectively address this challenge, multi-radar collaborative detection technology has emerged. In a multi-radar collaborative detection network, target allocation is one of the core tasks of system resource management. The purpose of target allocation is to allocate tracking targets to each radar node before the target tracking task, considering the influence of detection performance, system resources, collaborative efficiency, geographical environment, etc., so as to improve the accuracy, efficiency, and stability of subsequent target tracking tasks. However, the multi-radar collaborative detection target allocation problem is a typical dynamic optimization problem, and its complexity is mainly reflected in aspects such as the high maneuverability of targets, the finiteness of resources, and the uncertainty of the environment. Therefore, it is a difficult problem to model and solve the decision-making for the target allocation problem in a multi-radar multi-target scenario. Summary of the Invention
[0003] The purpose of the present invention is to provide a multi-radar collaborative detection target allocation method for multi-targets, aiming at the deficiencies of the prior art. By predefining a combined action pool and collaborative optimization of a two-layer policy, it solves the complexity problem of the high-dimensional discrete action space and explicitly integrates the engineering constraints of the radar system to improve the target tracking efficiency and resource utilization rate. To achieve the above purpose, the technical solutions adopted by the present invention are as follows:
[0004] Step 1: For the multi-radar collaborative detection multi-target scenario, the present invention combines the background knowledge of multi-radar collaborative detection, deeply studies multiple factors that need to be considered in the collaborative detection process, and proposes detection range constraints, target visibility constraints, tracking target quantity constraints, equipment switching times constraints, equipment usage time constraints, and collaborative detection coverage range constraints. The detection performance of the multi-radar system is evaluated through these six constraint conditions.
[0005] Step 2: According to the six constraint conditions proposed for the multi-radar collaborative detection scenario in Step 1, an evaluation model for target matching degree is constructed based on the AHP method. The target layer of this model is set as the multi-radar collaborative detection target matching degree, the criterion layer is set as B1 detection performance, B2 system efficiency, and B3 collaborative efficiency. The criterion layer B1 includes indicators C1 detection range and C2 target visibility. The criterion layer B2 includes indicators C3 number of tracked targets, C4 number of device switches, and C5 device usage time. The criterion layer B3 includes indicator C6 collaborative detection range. The present invention uses the AHP method to analyze the importance of the six indicators and obtains the corresponding weights γ1, γ2, γ3, γ4, γ5, γ6 of each indicator, and evaluates the matching degree of the target through the six indicators C1 to C6.
[0006] Step 3: According to the six indicator weights obtained in Step 2, a Markov decision process model is built for the target allocation problem based on the target matching degree evaluation model. The state space is set as: where where respectively represent the relative distance and visibility between radar node m and N targets, represents the cumulative usage time of the radar node at time t, represents the action of tracking targets of radar node m at time t - 1. The action space is set as A t ={a1, a2,..., a n}, a i ={c1, c2,..., c n}, which represents the tracked targets selected by each radar node from the combined action pool. Among them, c i is a binary number, 0 means not selected, 1 means selected. The combined action pool is an optional action pool generated according to the total number of targets and the maximum number of targets that a single radar node can track. Using the method of the combined action pool can effectively balance the dimension of the action space and the exploration difficulty.
[0007] Transition probability formula: P(S t+1 |S t , A t ) = P m ·P v ·P r , P m represents the target movement probability model. The target movement model adopts a random model, which specifically means moving a distance of speed v in a random direction at each moment. P v represents the target visibility probability model. This model adopts a random model, which specifically means that there is a certain target that is invisible to the three radar nodes every h steps. P rIt represents the resource consumption model, specifically the cumulative usage time of each radar node in the multi-radar system at that moment. Improving the randomness of the transfer probability can effectively improve the robustness of the detection system.
[0008] Reward function: It represents the weighted sum of rewards and corresponding penalties obtained by all radar nodes for the six performance indicators mentioned in Step 2 at time t, where γ k represents the weight of performance indicator k, represents the reward obtained by radar node i for performance indicator k, represents the penalty for each radar adopting an invalid action, represents the target loss penalty.
[0009] Step 4: According to the Markov decision process model obtained in Step 3, a target matching decision-solving framework is proposed based on the hierarchical multi-agent deep reinforcement learning algorithm. The hierarchical multi-agent deep reinforcement learning algorithm is divided into an outer layer policy (target selection layer) and an inner layer policy (action optimization layer). Combining with the predefined action pool, it significantly reduces the complexity of the high-dimensional discrete state space and action space. Among them, the input of the outer layer policy is the state space S t , and the output is the candidate target actions that meet the performance indicators C1 to C5 Its expression is the same as the action space A t , the input of the inner layer policy is the candidate target actions and the output is the tracking target actions that meet the performance indicator C6 The expression is the same as the action space A t . Through the training and optimization of the double-layer policy, the final target allocation policy is obtained. This allocation policy can input the state space information and output the tracking target combination with the highest matching degree for each radar node.
[0010] The beneficial effects of the present invention are as follows:
[0011] Aiming at the target allocation scenario of multi-radar and multi-flying target cooperative detection, the present invention comprehensively considers three aspects of detection performance, system effectiveness, and cooperation efficiency to construct a target matching degree evaluation model. This model fits the actual scenario and has dynamic adaptability.
[0012] The method proposed in the present invention, through the combination of predefined action pool and double-layer policy collaborative optimization, can effectively solve the problem of non-convergence in training caused by high-dimensional state space and high-dimensional action space in complex problems. Brief Description of the Drawings
[0013] Figure 1 is the flowchart of the solution of the present invention;
[0014] Figure 2 is the schematic diagram of the multi-radar system detecting multiple targets;
[0015] Figure 3 Schematic diagram of the target matching degree evaluation model based on AHP;
[0016] Figure 4 Algorithm flowchart of the present invention;
[0017] Figure 5 Comparison chart of the results of the algorithm of the present invention and other algorithms for ten radar nodes and three flying target scenarios;
[0018] Figure 6 Comparison chart of the results of the algorithm of the present invention and other algorithms for ten radar nodes and five flying target scenarios;
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Specific embodiments
[0020] The following will further elaborate on the specific implementation steps of the present invention in conjunction with the drawings.
[0021] The present invention proceeds according to the process in the attached Figure 1 and.
[0022] Step 1: As shown in the attached Figure 2 , the figure uses four radar nodes and four targets to illustrate the multi-radar system collaborative detection of multi-target scenarios. In combination with the background of multi-radar collaborative detection, the present invention proposes multiple constraint conditions:
[0023] Constraint condition 1: Detection range constraint. In a multi-radar system, each radar node has its own detection coverage range. In this detection range, the closer the target is, the less resources the radar node consumes to track the target, and the higher the priority of the target;
[0024] Constraint condition 2: In actual detection scenarios, each radar node is often affected by interference, weather, and terrain environment, resulting in reduced visibility of a certain flying target, thereby reducing the tracking efficiency of the radar node for this target;
[0025] Constraint condition 3: At the same time, a radar node transmitting multiple beams to track multiple targets will increase resource consumption. Therefore, it is necessary to avoid a single radar node tracking multiple flying targets simultaneously;
[0026] Constraint condition 4: A radar node switching the tracking beam to switch the tracking target will cause waste of system resources. Therefore, it is necessary to avoid frequent target switching;
[0027] Constraint 5: Prolonged use of a certain radar node will cause excessive consumption of the resources of that radar node. Therefore, it is necessary to balance the usage time of each radar node;
[0028] Constraint 6: Cooperative detection can greatly improve the detection ability of the radar system. Multiple radar nodes detecting the same target can consume fewer system resources to achieve better detection performance.
[0029] The present invention evaluates the detection performance, system effectiveness, and cooperation efficiency of the multi-radar cooperative detection system according to the above six constraints.
[0030] Step 2: As shown in the appendix Figure 2 and Figure 3 shown, according to the six constraints proposed for the multi-radar cooperative detection scenario in Step 1, the present invention constructs a target matching degree evaluation model based on the Analytic Hierarchy Process (AHP). First, set the target layer as the target matching degree evaluation model, and set three criteria in the criterion layer: B1 Detection Performance Criterion, B2 System Efficiency Criterion, B3 Cooperation Efficiency Criterion. Under the three criteria, set several indicators respectively: C1 Detection Distance Indicator, C2 Target Visibility Indicator, C3 Number of Tracked Targets Indicator, C4 Beam Switching Times Indicator, C5 Usage Time Indicator, C6 Cooperative Detection Coverage Range Indicator; Secondly, construct the judgment matrix, its maximum eigenvalue, and eigenvector for each criterion layer corresponding to the indicator layer respectively. The parameters of each judgment matrix are shown in Tables 1-3. The RI values set for the judgment matrices of order 1-9 by the Analytic Hierarchy Process are shown in Table 4. The Analytic Hierarchy Process defines the consistency test index as: and where λ max is the maximum eigenvalue of the judgment matrix, and n represents the order of the judgment matrix. When CR < 0.1, it means that the degree of inconsistency of A is within the allowable range. Calculate that the CR values of the judgment matrices constructed by the present invention are all less than 0.1, meeting the consistency test. Finally, synthesize the weights according to the eigenvectors of each judgment matrix to obtain the weights of the six judgment indicators C1, C2, C3, C4, C5, and C6 as: 0.27, 0.27, 0.11, 0.11, 0.06, 0.18.
[0031] Table 1 A-B Judgment Matrix
[0032]
[0033] Table 2 B1-C
[0034]
[0035] Table 3 B2-C Judgment Matrix
[0036]
[0037] Average Random Consistency Index RI Values of Judgment Matrices at All Levels in Table 4
[0038] n 1 2 3 4 5 6 7 8 9 RI 0 0 0.58 0.90 1.12 1.24 1.32 1.41 1.45
[0039] Step 3: According to the weights of the six indicators obtained in Step 2, use the Markov decision process method to model the target matching problem. Construct the state space as follows: where represents the relative distance between radar node m and N targets at time t, and the relative distance calculation formula is: represents the ratio of the distance between the coordinates of radar node m and target n to the maximum detection range of radar node m, where (x m , y m , z m ) are the coordinates of the radar node, is the maximum detection range of radar node m, (x n , y n , z n ) are the coordinates of the target. When the distance between the radar and the target is closer, is larger, and the detection effect is better. This expression reflects the C1 performance index, represents the visibility between radar node m and N targets at time t. In the present invention, the visibility is quantified as shown in the formula:
[0040]
[0041] Through this expression, the C2 performance index is reflected, represents the usage duration of radar node m at time t, reflecting the C5 performance index, represents the target selection action of radar node m at time t - 1, which can provide a judgment basis for the C4 performance index. Construct the action space as: A t ={a1, a2,..., a n}, a i ={c1, c2,..., c n}, where c i is a binary number, 0 means selected, and 1 means not selected. The action space A tIndicates the tracking targets selected by each radar node from the combined action pool. The combined action pool is an optional action pool generated based on the total number of targets and the maximum tracking number of a single radar node. Using the method of the combined action pool can effectively balance the dimensionality of the action space and the exploration difficulty. Construction rules of the action pool: (1) For any k ∈ [0, maximum number of transmitting beams of a single radar node], generate all combinations of selecting k targets from the total number of targets N. (2) Each combination is represented as a binary vector of length N, where the index represents the target number, and 1 indicates the selected target while 0 indicates the unselected target. The constructed state transition probability formula is: P(S t+1 |S t ,A t ) = P m ·P v ·P r , P m represents the target movement probability model. The target movement model adopts a random model, specifically represented as moving a distance of speed v in a random direction at each moment. P v represents the target visibility probability model. This model adopts a random model, specifically represented as a certain target being invisible to the three radar nodes every 20 steps. P r represents the resource consumption model, specifically represented as the cumulative usage time of each radar node in the multi-radar system at this moment. By increasing the randomness of the transition probability, the robustness of the radar detection system can be improved. Construction of the reward function: represents the sum of rewards and corresponding penalties obtained by all radar nodes for six performance indicators at time t, represents the reward obtained for the detection range indicator. I(*) represents judging whether the condition in the parentheses holds, with 1 for true and 0 for false, represents whether radar node i tracks target j, represents the reward for the visibility indicator. N represents the total number of targets, represents the reward for the number of tracked targets indicator, where represents the action vector represents the number of binary 1s in represents the reward obtained for the number of device switches, represents the reward obtained for the usage time indicator, where represents the cumulative usage time of radar node i at time t, represents the reward obtained for the collaborative detection coverage range indicator, C j represents the target combination in which the number of radar nodes tracking the same target as radar node i is not less than 3 among the targets tracked by radar node i. |C j | represents the number of targets in the combination, represents the mean of the distances between pairs of multiple radars that constitute collaborative detection. represents the number of radar nodes that simultaneously detect target j at time t. represents the relative distance between radar node i and radar node j. represents the distance covariance between multiple radars that constitute collaborative detection.
[0042] represents the penalty for the absence of a target in radar tracking. represents the penalty for the existence of a target without being tracked by a radar node;
[0043] Step 4: According to the Markov decision process model obtained in Step 3, a target matching decision-solving framework is proposed based on the hierarchical multi-agent deep reinforcement learning algorithm. Due to the complexity of the multi-radar multi-target collaborative detection target matching problem addressed by the present invention, the dimensions of the state space and action space in the MDP model are too large, and using typical deep reinforcement learning algorithms will result in difficult convergence and instability. Therefore, the present invention proposes a hierarchical multi-agent deep reinforcement learning algorithm, as shown in the appendix Figure 4 This algorithm uses a two-layer MADDPG network. The outer-layer MADDPG algorithm contains M agents, each agent corresponding to a radar node. The input of each agent is the local observation of that radar node, that is The output is the action of initially screening targets. The inner-layer MADDPG algorithm contains M agents, each agent corresponding to a radar node. The input of each agent is the initially screened actions output by all the agents in the outer layer, and the output is the optimized action. In this algorithm, the outer-layer MADDPG algorithm mainly considers five performance indicators C1 to C5 to initially screen out the targets that are relatively matched with each radar node. The inner-layer MADDPG algorithm considers the C6 performance indicator to obtain the final matching targets.
[0044] The specific process of the two-layer MADDPG algorithm is as follows:
[0045] (1) Initialize the environment, initialize the Actor network and Critic network of each agent, and generate a combined action pool;
[0046] (2) Loop and execute steps 3 to 11, and increment the number of training rounds by one each time the loop is executed until the maximum number of training rounds is reached;
[0047] (3) Reset the environment and obtain the initial state;
[0048] (4) Loop and execute steps 5 to 8 until the episode ends;
[0049] (5) Input the initial state, and each actor network of the outer-layer MADDPG generates an outer-layer action;
[0050] (6) Input the outer-layer action, and each actor network of the inner-layer MADDPG generates the inner-layer action;
[0051] (7) Input the inner-layer action into the action pool to obtain the target selection action of each radar;
[0052] (8) Execute the action and update the environment;
[0053] (9) The inner-layer action obtains the C6 index reward, and stores the state, action, new state, and reward in the experience pool;
[0054] (10) The outer-layer action obtains the weighted reward of the six indexes C1 to C6, and stores the state, action, new state, and reward in the experience pool;
[0055] (11) If the number of the experience pool reaches the sampling number, the two-layer MADDPG algorithm samples data from their respective experience pools for network update.
[0056] Step 5: Based on the gym framework, build a multi-radar multi-target cooperative detection target matching scenario with 10 radar nodes and a maximum of 3 targets. The radar coordinates are fixed, and the coordinate range is in the interval [50, 450]. The maximum detection range of each radar is 300, and the target coordinate range is in the interval [0, 500]. Each target flies randomly in one direction at a speed of v = 5 per round, and one target is randomly set to be invisible to the three radar nodes every 20 steps. Select the algorithm of the present invention, the MADDPG algorithm, the DQN algorithm, the DoubleDQN algorithm, and the DuelingDQN algorithm for comparison. The parameters of the algorithm of the present invention are shown in Table 5, and the comparison results are as attached Figure 5 as shown.
[0057] It can be Figure 5 seen that the algorithm proposed by the present invention converges at 2500 rounds, and the fluctuation of the reward after stabilization is very small, while there are still large fluctuations in the other comparison algorithms after tending to converge, and the reward convergence value of the algorithm proposed by the present invention is also significantly higher than that of the comparison algorithms.
[0058] Table 5 List of parameters of the hierarchical MADDPG algorithm
[0059] Number of training rounds 20000 Experience pool size 500000 Sampling data batch size 1024 Soft update parameter 0.005 Loss factor 0.99 Actor network learning rate 3e-4 Critic network learning rate 3e-4 Actor network structure [[input_dim, 128], [128, 64], [64, output_dim]] Critic network structure [[input_dim, 128], [128, 64], [64, 1]] Activation function Leak_relu
[0060] Step 6: On the basis of the experiment in Step 5, increase the number of flying targets to 5, and the comparison results are as Figure 6 shown.
[0061] It can be Figure 6It can be seen that when the number of targets is increased to 5, the capacity of the combined action pool increases and the dimension of the action space increases, resulting in large fluctuations in the early stage of the training of the algorithm of the present invention. However, the algorithm converges at about 3000 rounds. The convergence reward value is comparable to that of the MADDPG algorithm and higher than that of the DQN series algorithms. Moreover, the algorithm proposed in the present invention is stronger than other algorithms in terms of stability.
Claims
1. A multi-target multi-radar cooperative detection target allocation method, characterized in that: The following steps are involved: Step 1: For the scenario of multi-radar collaborative detection of multiple targets, the detection performance of the multi-radar system is evaluated through six constraints; Step 2: According to the six constraints, a target matching evaluation model is constructed based on the AHP method, and the six indicators and corresponding weights included in the target matching evaluation criterion layer are obtained; Step 3: According to the corresponding weights of the six indicators, the Markov decision process modeling of the target allocation problem is carried out based on the target matching evaluation model; Step 4: According to the Markov decision process model, a target matching decision solution framework is proposed based on the hierarchical multi-agent deep reinforcement learning algorithm to output the tracking target combination with the highest matching degree for each radar node.
2. The multi-target multi-radar cooperative detection target allocation method according to claim 1 is characterized in that: The six constraints described in step 1 are specifically as follows: for the multi-radar collaborative detection of multiple targets scenario, combined with the background knowledge of multi-radar collaborative detection, detection range constraints, target visibility constraints, tracking target quantity constraints, device switching times constraints, device usage time constraints and collaborative detection coverage constraints are proposed.
3. The multi-target multi-radar cooperative detection target allocation method according to claim 2 is characterized in that: The step 2 is specifically implemented as follows: according to the six constraints, a target matching evaluation model is constructed based on the AHP method; the target layer of the model is set to the multi-radar collaborative detection target matching, the criterion layer is set to B1 detection performance, B2 system efficiency, B3 collaborative efficiency, the criterion layer B1 includes the indicators C1 detection range and C2 target visibility, the criterion layer B2 includes the indicators C3 number of tracked targets, C4 number of device switching times, C5 device usage time, and the criterion layer B3 includes the indicator C6 collaborative detection range; the AHP method is used to analyze the importance of the six indicators and obtain the weights γ1, γ2, γ3, γ4, γ5, and γ6 corresponding to each indicator, and the matching degree of the target is evaluated by the six indicators C1 to C6.
4. The multi-radar cooperative detection target allocation method for multiple targets according to claim 3 is characterized in that: The specific process of Markov decision process modeling is: Set up the state space as: in in They represent the relative distance and visibility between radar node m and N targets respectively, represents the cumulative usage time of the radar node at time t, Represents the target tracking action of radar node m at time t-1; set the action space as A t ={a1,a2,...,a n },a i ={c1,c2,...,c n }, represents the tracking target selected by each radar node from the combined action pool, where c i It is a binary number, 0 means no selection, 1 means selection, and the combined action pool is an optional action pool generated according to the total number of targets and the maximum number of tracking of a single radar node; Transition probability formula: P(S t+1 |S t ,A t )=P m ·P v ·P r , where P m represents the target movement probability model. The target movement model adopts a random model, which is expressed as the distance of random movement in one direction at each moment with a speed v. P v represents the target visibility probability model, which adopts a random model. Specifically, every h steps, there is a target that is invisible to the three radar nodes, P r It represents the resource consumption model, which is specifically expressed as the accumulated usage time of each radar node in the multi-radar system at that moment; Reward function: represents the weighted sum of rewards and corresponding penalties obtained by all radar nodes at time t for the six indicators mentioned in step 2, where γ k represents the weight of performance indicator k, represents the reward obtained by radar node i for performance indicator k, Indicates the penalty for each radar taking invalid actions, Indicates the target loss penalty.
5. The multi-target coordinated detection target allocation method for multiple radars according to claim 4 is characterized in that: The step 4 is specifically implemented as follows: the hierarchical multi-agent deep reinforcement learning algorithm is divided into an outer strategy and an inner strategy, combined with a predefined action pool, where the input of the outer strategy is the state space S t , the output is the candidate target action that meets the performance indicators C1 to C5 Its expression is consistent with the action space A t Consistent, the input of the inner strategy is the candidate target action The output is the tracking target action that meets the C6 performance index Expression and action space A t The final target allocation strategy is obtained by training and optimizing the double-layer strategy, which can input state space information and output the tracking target combination with the highest matching degree of each radar node.