A Multi-UAV Cooperative Target Search Method for Urban Environments Based on NRO-QMIX

Through the NRO-QMIX-based method, a multi-drone search environment architecture is built, target detection rules and reward functions are designed, regional information trusted maps are established, and the drone execution actions are optimized. The inefficiency of multi-drone collaborative target search and information collection in urban environments is solved, and efficient task completion is achieved.

CN119902564BActive Publication Date: 2025-07-04THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510388161.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-04
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

In urban environments, multi-UAV coordinated target search and information collection tasks face complexity and uncertainty, and existing methods are difficult to effectively deal with unknown areas, resulting in inefficient searches or incomplete information collection.

Method used

Using the NRO-QMIX-based method, by constructing a multi-drone search environment architecture, designing drone search target detection rules and reward functions, establishing a trusted map of regional information, and designing an improved NRO-QMIX algorithm to optimize the execution actions of each drone, realize the mapping of global and local actions, and enhance the environmental exploration capabilities and algorithm convergence speed.

Benefits of technology

The efficiency of multi-drone collaborative target search and information collection tasks has been improved, and the problem of difficult combination optimization tasks in complex urban environments is solved. The environmental exploration capabilities of drones have been enhanced and the algorithm convergence speed has been accelerated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119902564B_ABST
    Figure CN119902564B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-UAV collaborative target search method based on NRO-QMIX for urban environments, which includes constructing a multi-UAV search environment architecture based on the unknown area to be searched, designing UAV search target detection rules and reward functions, establishing a regional information credibility graph that maps the unknown area, and proposing to fuse target search and information collection to achieve the optimization problem of maximizing the update of the regional credibility graph. An improved NRO-QMIX algorithm is designed to solve the optimization problem, obtain the optimal solution of the execution actions of each UAV, and complete multi-UAV collaborative target search and information collection in complex urban environments. The present invention solves the problem that it is difficult to simultaneously perform the combined optimization tasks of target search and information collection in complex urban environments, and the problem that the existing reinforcement learning methods use greedy strategies to select actions, resulting in low exploration efficiency of the agent and the environment, and realizes multi-UAV collaborative target search and information collection in urban environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) target search, and particularly to a multi-UAV cooperative target search method based on NRO-QMIX in an urban environment. Background Art

[0002] In the face of the complexity and uncertainty of the urban environment, multi-UAV cooperative search technology plays an important role in target search and area information collection tasks. Due to the complex and changeable scenarios, the current solution methods for multi-UAV cooperative search are roughly divided into two categories: rule-based search methods and mathematical programming-based search methods.

[0003] Rule-based search methods are usually used for large-scale coverage search tasks. This method pays more attention to the problems of missed search areas and resource consumption existing in existing coverage algorithms, and is committed to optimizing the search formation. Although it can effectively solve the defects of traditional coverage algorithms and perform well in coverage search environments and fixed UAV formations, when the environment is complex or the number of people changes, it is necessary to re-plan the search paths of UAVs.

[0004] Mathematical programming-based methods are used for more complex environments. They combine optimization theory to model the multi-UAV cooperative target search (MCTS) problem as an optimization problem with the minimum search time or the maximum coverage rate as the objective function, and usually use swarm intelligence algorithms for search and solution. However, the time and space complexity of this method increase rapidly with the increase in the number of UAVs, resulting in unsatisfactory results.

[0005] With the rise of reinforcement learning, solving the MCTS problem based on reinforcement learning has become widely popular and has become a new paradigm for solving the problems of multi-target cooperative search and information collection in cities. However, in the urban environment, the MCTS task faces a high degree of uncertainty in unknown areas, and complex terrains and obstacles significantly increase the task difficulty. Existing methods often cannot effectively handle uncertainties, resulting in low search efficiency or incomplete information collection; in the MCTS task, the complexity of the urban environment requires UAVs to quickly adapt to new situations and discover optimal strategies, while multi-agent reinforcement learning algorithms are less efficient in exploring new paths or strategies, resulting in slow convergence or poor effects of search tasks. Summary of the Invention

[0006] Object of the Invention: The object of the present invention is to provide a multi-UAV cooperative target search method based on NRO-QMIX in an urban environment, so as to improve the efficiency of multi-UAV cooperative target search and information collection tasks in an unknown urban environment.

[0007] Technical solution: To achieve the above object, a multi-UAV collaborative target search method for urban environment based on NRO-QMIX of the present invention includes the following steps:

[0008] Step 1: Based on the area to be searched Construct a multi-UAV search environment architecture;

[0009] Step 2: Design UAV search target detection rules and reward functions;

[0010] Step 3: Establish a regional information credibility graph of the mapped area and propose an optimization problem that integrates target search and information collection to maximize the update of the regional credibility graph;

[0011] Step 4: Design an improved NRO-QMIX algorithm to solve the optimization problem, obtain the optimal solution of the execution actions of each UAV, and complete multi-UAV collaborative target search and information collection in the urban environment;

[0012] The improved NRO-QMIX algorithm includes separately processing the observation information of each UAV to generate a local policy network and mixing the locals of all UAVs to generate a global mixing network. By constructing the constraint relationship between the global and the local , a mapping relationship between the global maximum value and the local maximum value is established to ensure that the local optimal action of each UAV is also the global optimal action; among them, the output layer of the policy network is the NROWAN noise network.

[0013] Among them, the multi-UAV search environment architecture described in Step 1 includes an environment model, a target and threat area model, and a UAV model;

[0014] Among them, the method for constructing the environment model is: rasterize the area to be searched and divide it into multiple squares, which respectively represent the number of rows and columns of the squares; define the squares as , and represent the center of the square as , where, , ;

[0015] The method for constructing the target and threat area model is: assume that each target only occupies one square, and there is only one target in each square. Define the event as whether there is a target in the square, indicates that there is a target in the current square, It indicates that there is no target in the current grid. In the unknown area, the preset probability of target existence in each grid is 0.5;

[0016] Define the threat area as a circle centered at the center position of a certain grid, which is expressed as:

[0017] ,

[0018] In the formula, is the center of the threat area, is the threat area number, is the radius of the threat area, R + represents a positive real number;

[0019] During the operation of the UAV, it needs to maintain a certain distance from the threat area, which is expressed as:

[0020] ,

[0021] In the formula, is the UAV coordinates, is the preset safety distance between the UAV and the dangerous area;

[0022] The method for constructing the UAV model is as follows: The UAV performs particle motion, and the kinematics of a single UAV is expressed as:

[0023] ,

[0024] In the formula, is the speed, is the heading angle, is the angular velocity;

[0025] The multi-UAV system composed of single UAVs is expressed as:

[0026] ,

[0027] In the formula, is the total number of UAVs;

[0028] Assume that the UAV moves only one grid at each moment. The mathematical relationship of anti-collision positions between multiple UAVs is expressed as:

[0029] ,

[0030] In the formula, is the preset safety distance for anti-collision between multiple UAVs, 、 is the UAV 、j coordinates.

[0031] Among them, the design method of the UAV search target detection rule described in step 2 is:

[0032] Let event represent that at time t, the UAV detects a target in the grid . Let represent that at time t, the UAV does not detect a target in the grid . Then the probability that the UAV detects a target in the grid

[0033] is expressed as:

[0034] wherein, is the detection rate, is the false alarm rate, is the missed detection probability, is the true missed detection probability, represents that there is a target in the grid , represents that there is no target in the grid ;

[0035] The method for updating the probability of the existence of a target in the grid is as follows:

[0036] ,

[0037] wherein, is the probability of the UAV detecting the existence of a target in the grid at time t;

[0038] The UAV judges according to the probability value of the existence of a target in the current grid to determine the value of the event :

[0039] ,

[0040] When the probability of the existence of a target in the grid exceeds the threshold , the UAV judges that there is a target in the grid .

[0041] Among them, the reward function described in step 2 includes target reward, information collection reward, obstacle avoidance penalty, and time step penalty;

[0042] Among them, the target reward represents the reward given when the UAV searches for a new target, and the UAV is only rewarded when each target is first discovered. The target reward function is expressed as:

[0043] ,

[0044] Among them, is the coefficient for adjusting the size of the target reward, is the total number of drones, is the different number of the drones, is the drone at time t in the grid the probability of detecting the presence of the target; is the threshold of the probability of the presence of the target;

[0045] The information collection reward means that within a limited time, the drones fly to cover as many grids as possible. The information collection reward function is expressed as:

[0046] ,

[0047] In the formula, is the coefficient for adjusting the size of the information collection reward, is the number of times the grid is searched by the drones at time t;

[0048] The obstacle avoidance penalty means to avoid collisions between drones during the mission. The obstacle avoidance penalty function is expressed as:

[0049] ,

[0050] In the formula, is the coefficient for adjusting the penalty size of the drone collision danger zone, is the coefficient for adjusting the penalty size of the collision between drones; represents the threat area, s is the threat area number, , are the coordinates of the drones , j, is the preset safe distance between the drone and the danger zone, is the preset safe distance for anti-collision between multiple drones;

[0051] The time step penalty means that the flight time of the drones is fixed. The time step penalty function is:

[0052] ,

[0053] In the formula, is the coefficient for adjusting the penalty size exceeding the time step;

[0054] The reward obtained by all drones is the total reward. The total reward function is expressed as:

[0055] .

[0056] Among them, the method for establishing the regional information credibility graph described in step 3 is: establish a mapping area of the regional information graph , where each grid represents the regional information credibility of the corresponding square . The regional credibility graph is expressed as:

[0057] ,

[0058] In the formula, is the regional credibility graph at time t. If the square belongs to the threat area , then the regional credibility graph is set to 0;

[0059] is the change in the regional credibility graph after the UAV information collection at time t, and is expressed as:

[0060] ,

[0061] is the number of times the square has been searched at time t, is the initial regional information credibility of the square ;

[0062] is the change in the regional credibility graph after the UAV searches for the target at time t, and is expressed as:

[0063] ,

[0064] In the formula, the event is that the UAV detects the target in the square , represents that the UAV does not detect the target in the square , is the probability of detecting the existence of the target by the UAV in the square at time t.

[0065] Among them, the optimization problem described in step 3 is expressed as:

[0066] ,

[0067] In the formula, is the regional credibility graph at time t, , represent the horizontal and vertical coordinate ranges of the UAV , represents the UAV in the area Fly inside.

[0068] Among them, the data processing process of the improved NRO-QMIX algorithm described in step 4 is as follows:

[0069] The UAV Current observation value And the action at the previous moment Are input into the policy network of each agent network for processing, and the individual action value function is output , that is, local ; Among them, Indicates the UAV Its own historical actions and observation values, expressed as: , Indicates the UAV predicted by the policy network Execute the action at the current moment;

[0070] Each UAV Of the local And the global observation s are input into the mixing network to generate the joint action value function, that is, the global , expressed as:

[0071] ;

[0072] Establish the mapping relationship between the global Maximum value and the local Maximum value is expressed as:

[0073] ,

[0074] The constraint relationship to satisfy the mapping relationship is expressed as:

[0075] ,

[0076] In the formula, Is the mixing network parameter, Is the total number of UAVs.

[0077] Among them, the network structure of the NROWAN noise layer described in step 4 is expressed as:

[0078] ,

[0079] In the formula, Is the learnable parameter, Is the noise.

[0080] Among them, the process of solving the optimization problem by the improved NRO-QMIX algorithm is as follows:

[0081] (1) At each discrete moment, the UAV All interact with the environment in the exploration area. At time t, the UAV generates its own observation value at the current moment , and for each UAV its own observation value and the action at the previous moment as well as the global observation are input into the NRO-QMIX neural network, and the actions with the largest global and local values for each UAV are output . The UAV executes this action and obtains the total reward of the multi-UAV at the current moment t , thus entering the next moment and completing a complete interaction loop;

[0082] (2) The UAV starts a new loop, continues to interact with the environment, and continuously accumulates rewards. When a limited number of moments have passed, all loops end, and the overall of the multi-UAV is obtained. During several rounds of training, the NRO-QMIX neural network is continuously updated according to the previous training experience and the designed loss function until the training is completed;

[0083] (3) Using the trained NRO-QMIX neural network, the optimal action executed by the UAV under the current observation value is obtained.

[0084] Among them, the loss function is expressed as:

[0085] ,

[0086] In the formula, are the parameters of the target network, reflects the noise level, is the proportionality coefficient, is the discount coefficient, is the total reward of the multi-UAV at time t, is the noise value function of the target global , is the global noise value function;

[0087] Among them, , ,

[0088] In the formula, is the input dimension of the last layer of the NROWAN noise network, is the number of output actions, is​ The final expected value, is the reward at the current time t, is the maximum reward, is the minimum reward, is the noise variance of the weight, is the noise variance of the bias.

[0089] Beneficial effects: The present invention has the following advantages: 1. By introducing a region information credible map that integrates the target search task and the information collection task, the method of the present invention transforms the combinatorial optimization problem into a deterministic task that only needs to optimize the region information credible map, and solves the problem that it is difficult to simultaneously perform the combinatorial optimization task of target search and information collection in a complex urban environment;

[0090] 2. The method of the present invention is improved based on the existing QMIX algorithm, incorporates the NROWAN noise network, and optimizes the reward function based on task requirements and environmental characteristics, which can effectively enhance the environmental exploration ability of the unmanned aerial vehicle and accelerate the convergence speed of the algorithm, and solves the problem that the existing reinforcement learning method uses a greedy strategy to select actions, resulting in low exploration efficiency of the agent and the environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] Figure 1 is the flowchart of the multi - unmanned - aerial - vehicle collaborative target search method described in the present invention;

[0092] Figure 2 Schematic diagram of environmental modeling;

[0093] Figure 3 is the architecture diagram of the improved NRO - QMIX algorithm;

[0094] Figure 4 is the schematic diagram of the data processing code of the improved NRO - QMIX algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0095] The technical solutions of the present invention will be described in detail below in conjunction with the embodiments and the drawings.

[0096] As Figure 1 shown, a multi - unmanned - aerial - vehicle collaborative target search method based on NRO - QMIX described in the present invention includes the following steps:

[0097] I. Construct a multi - unmanned - aerial - vehicle search environment architecture, including an environment model, a target and threat area model, and an unmanned - aerial - vehicle model, as Figure 2 shown.

[0098] (1) Construct an environment model

[0099] The area to be searched is rasterized and divided into multiple square grids, represent the number of rows and columns of the grid respectively; define the grid as , and represent the center of the grid as , where , .

[0100] (2) Construct the target and threat area model

[0101] Assume that each target only occupies one grid, and there is only one target in each grid; the description of whether there is a target in the unknown area adopts the Bernoulli distribution: define the event as whether there is a target in the grid, indicates that there is a target in the current grid, indicates that there is no target in the current grid. In the unknown area, the preset probability of the existence of a target in each grid is 0.5.

[0102] Define the threat area as a circle with the center position of a certain grid as the center, expressed as:

[0103] (1),

[0104] where is the center of the threat area, is the threat area number, is the radius of the threat area, R + represents a positive real number.

[0105] During the operation of the UAV, it needs to maintain a certain distance from the threat area, expressed as:

[0106] (2),

[0107] where is the UAV 's coordinates, is the preset safe distance between the UAV and the danger area.

[0108] (3) Construct the UAV model

[0109] The UAV performs particle motion, and the kinematics of a single UAV is expressed as:

[0110] (3),

[0111] where is the speed, is the heading angle, is the angular velocity.

[0112] The multi-UAV system composed of single UAVs is expressed as:

[0113] (4),

[0114] where, is the total number of drones.

[0115] Assume that each drone moves only one square at a time. Since each drone exists alone in a square at each moment t, when multiple drones cooperate in search, collision avoidance needs to be considered. The mathematical relationship of the collision avoidance positions between multiple drones is expressed as:

[0116] (5),

[0117] where, is the preset safe distance for collision avoidance between multiple drones, , are the coordinates of drone , j.

[0118] II. Design the target detection rules for drone search.

[0119] Based on the fact that the sensors carried by the drones themselves may have situations such as missed detection and false alarms, let event represent that at moment t, drone detects a target in square , represent that at moment t, drone does not detect a target in square . Then at moment t, drone does not detect a target at square . Combining the above event B, the following four possibilities will occur for different event combinations:

[0120] A. The sensor of drone detects a target at square at moment t, and there is a target in this square ;

[0121] B. The sensor of drone does not detect a target at square at moment t, but there is a target in this square ;

[0122] C. The sensor of drone detects a target at square at moment t, but there is no target in this square ;

[0123] D. The sensor of drone does not detect a target at square at moment t, and there is no target in this square ;

[0124] Unmanned aerial vehicle The probability of detecting a target at grid at time t is as shown in Equation (6):

[0125] (6),

[0126] where is the detection rate, is the false alarm rate, is the missed detection probability, is the true missed detection probability.

[0127] At each time t, according to Equation (6) and combined with Bayes' criterion, the update method of the target presence probability at time t is as follows:

[0128] (7),

[0129] where is the probability of detecting the presence of a target by the unmanned aerial vehicle in grid at time t.

[0130] The unmanned aerial vehicle judges according to the probability value of the presence of a target in the current grid to determine the event value to reduce situations such as missed detection and false alarm. The judgment formula is:

[0131] (8),

[0132] When the probability of the presence of a target in the grid exceeds the threshold , the unmanned aerial vehicle judges that there is a target in grid .

[0133] III. Based on the target detection rules for unmanned aerial vehicle search, establish a regional information credibility graph.

[0134] Establish a regional information graph of the mapped area , where each grid represents the credibility of the regional information corresponding to the grid .

[0135] The regional information credibility graph will change when the unmanned aerial vehicle searches for a target and collects information. The process of the unmanned aerial vehicle searching for a target and collecting information changes and progresses continuously with time t, and the regional information graph also changes accordingly. The change of the regional credibility graph after the unmanned aerial vehicle searches for a target is as follows:

[0136] ​(9),

[0137] Among them, is the change of the regional credibility graph after the UAV searches for the target at time t, is the grid Initial regional information credibility.

[0138] The change of the regional credibility graph after the UAV collects information is as follows:

[0139] (10),

[0140] Among them, is the change of the regional credibility graph after the UAV collects information at time t, is the number of times the grid is searched at time t.

[0141] (11),

[0142] Among them, is the regional credibility graph at time t. If the grid belongs to the threat area, the regional credibility graph is set to 0.

[0143] IV. Under the multi-UAV search environment architecture, an optimization problem is established to maximize the update of the regional credibility graph when each UAV searches for more targets, collects more information, and at each moment.

[0144] This optimization problem means that in a limited area, under the two constraints of avoiding collisions between multi-UAVs and avoiding threat areas, the multi-UAVs should complete the cooperative search task as excellently as possible.

[0145] (12),

[0146] Among them, 、 represent the horizontal and vertical coordinate ranges of the UAV , represents the UAV flies within the area .

[0147] V. Design the NRO-QMIX algorithm with centralized training and distributed execution to coordinate the search behaviors among multi-UAVs.

[0148] Among them, centralized training means training using global information. In the training stage, the algorithm considers the information of all UAVs to learn the global optimal strategy; distributed execution means relying on local information during execution. In the execution stage, each UAV makes decisions only based on the local information it observes, realizing distributed decision-making.

[0149] (1)Design a random noise mixed Q-value algorithm (NRO-QMIX), as Figure 3 shown.

[0150] In experiments, decomposed noise is usually used to generate noise for the input and output layers. The NRO-QMIX algorithm of the present invention improves the QMIX algorithm. Specifically, the MLP output layer of the agent network in the QMIX algorithm is replaced with an NROWAN (restricted random noise) noise layer for output, aiming to more effectively process the noise of the input and output layers, so as to better adapt to the high uncertainty and complexity in the urban environment and improve the efficiency and effect of multi-UAV collaborative search.

[0151] The NROWAN noise layer has high stability, enhancing the exploration strength of the agent while making the network training more stable. The network has a dimensional input and a dimensional output fully connected layer can be written as , and the network structure of the NROWAN noise layer can be written as:

[0152] (13),

[0153] where, are learnable parameters, and is noise.

[0154] The NRO-QMIX algorithm uses global information for training, but only relies on local information during execution. The specific implementation method for achieving efficient distributed decision-making in complex environments is as follows:

[0155] In the NRO-QMIX algorithm, an independent policy network is reserved for each UAV to process the observations of the UAV. . After passing through the NROWAN noise layer, the independent policy network outputs the individual action value function through the argmax operation, which is also called the local . The data processing process of the policy network is as follows:

[0156] The current observation of the UAV and the action at the previous moment are input into the policy network of each agent network i for processing, mining the useful information hidden in the historical observations, and the information is input into the NROWAN noise network for the argmax operation, and finally the individual action value function is output, which is also called the local . Among them, represents the historical actions and observations of the UAV itself,​ , represents the UAV predicted by the policy network to execute an action at the current moment.

[0157] Meanwhile, the NRO-QMIX algorithm retains a centralized mixing network, mixes the local of each UAV , and adds the global observation s (the global observation state is constructed from the global credible map and the overall probability map) to generate the joint action value function, also known as the global Q tot . That is, the local of each UAV and the global observation are input into the mixing network for mixing to generate the joint action value function, also known as the global Q tot .

[0158] (14),

[0159] where are the mixing network parameters.

[0160] By constructing the constraint relationship between the global Q tot and the local , the mapping relationship between the maximum value of the global Q tot and the maximum value of the local is established to ensure that the local optimal action of each UAV is also the global optimal action. The specific implementation process is as follows:

[0161] Explore the mapping relationship between the joint action value function in formula (14) and the local value function . By extracting the decentralized policy from the maximized joint value function, since taking argmax of the joint action value function is equivalent to taking argmax of each local action value function, while ensuring the same monotonicity, it is necessary to satisfy that taking the maximum value of Q tot and taking the maximum value of are sufficient. The mapping relationship is as follows:

[0162] (15),

[0163] To satisfy equation (14), the following constraint relationship needs to be satisfied:

[0164] (16),

[0165] where is the total number of UAVs.

[0166] The constraint relationship of formula (15) indicates that the global Qtot For each drone the local partial derivative must be non - negative. If the local of the drone increases, the global Q tot will not decrease. This constraint ensures that the local optimal action of each drone can also lead to the global optimal action, thus simplifying the multi - agent coordination problem.

[0167] (2)Design the loss function of the NRO - QMIX algorithm to achieve network update.

[0168] In the NRO - QMIX algorithm, a target network is adopted to maintain the stability of the training process and avoid oscillation and divergence. At the same time, the NRO - QMIX algorithm uses the noisy value function and to represent the global and the target global respectively. The NRO - QMIX loss function is as follows:

[0169] (17),

[0170] where, are the parameters of the target network, reflects the noise level, is the proportionality coefficient, is the discount coefficient, is the total reward of multiple drones at time t, is the noisy value function of the target global , is the global noisy value function.

[0171] Among them: (18),

[0172] (19),

[0173] In equation (18), is the input dimension of the last layer of the NROWAN noise network, is the number of output actions.

[0174] In equation (19), is the final expected value of , is the reward at the current time t, is the maximum reward, is the minimum reward, is the noise variance of the weight, is the noise variance of the bias.

[0175] 6. Design the reward function for the multi-UAV collaborative target search and information collection task.

[0176] The reward function is designed to guide the multi-UAVs to search for as many targets as possible, collect as much area information as possible, and avoid potential risk areas in the area. The present invention designs the following four rewards in combination with the actual targets:

[0177] (1) Design the target reward based on the human-machine search target detection rule

[0178] When a UAV searches for a new target, a large reward is given. Each target will only give the UAV a reward when it is first discovered. The target reward function is:

[0179] (20),

[0180] where, is the coefficient for adjusting the size of the target reward, is the total number of UAVs, is the different number of the UAV, .

[0181] (2) Information collection reward

[0182] The information collection reward is designed to let the UAVs fly over as many squares as possible within a limited time. The information collection reward function is:

[0183] (21),

[0184] where, is the coefficient for adjusting the size of the information collection reward, is the number of times the square is visited by the UAV at time t.

[0185] (3) Obstacle avoidance penalty

[0186] The obstacle avoidance penalty is designed to prevent the UAVs from colliding with each other during the task. The obstacle avoidance penalty function is:

[0187] (22),

[0188] where, is the coefficient for adjusting the penalty for the UAV collision danger zone, is the coefficient for adjusting the penalty for the UAVs colliding with each other.

[0189] (4) Time step penalty

[0190] The design of the time-step penalty is due to the limited energy carried by the UAV and the fixed flight time. The time-step penalty function is as follows:

[0191] (23),

[0192] where is the coefficient for adjusting the size of the penalty beyond the time step.

[0193] During the process of the reinforcement learning task, the total reward obtained by all UAVs is accumulated as the total reward, and the total reward is:

[0194] (24).

[0195] VII. Use the NRO-QMIX algorithm to solve the optimization problem and obtain the optimal solution for the execution actions of each UAV, as Figure 4 shown.

[0196] At each discrete moment, the UAV interacts with the environment in the exploration area. At time t, the UAV generates its own observation value at the current moment ;

[0197] Input the own observation value of each UAV and the action at the previous moment as well as the global observation into the NRO-QMIX neural network, and output the action corresponding to the global and local with the maximum value . The UAV executes this action , and the total reward of the multi-UAVs at the current moment t .

[0198] The environment changes with the actions of the UAVs, thus entering the next moment and completing a full interaction cycle. Subsequently, the UAVs start a new cycle and continue to interact with the environment. In this way, the UAVs keep taking actions and interacting with the environment, and continuously accumulate rewards. When a limited number of moments have passed, all cycles end, and this round of the task session ends, obtaining the overall of the multi-UAVs.

[0199] During several rounds of training, the NRO-QMIX algorithm is continuously updated according to the previous training experience and the designed loss function until the training is completed.

[0200] Using the trained NRO-QMIX neural network, obtain the UAV The optimal action executed under the current observation value enables multiple UAVs to maximize the update of the regional trust map within the least number of time instants to complete the collaborative target search of multiple UAVs in a complex urban environment, and at the same time, the maximum cumulative reward can also be obtained.

Claims

1. A multi-UAV collaborative target search method for urban environment based on NRO-QMIX, characterized in that It includes the following steps: Step 1: Based on the area to be searched Construct a multi-UAV search environment architecture; Step 2: Design the target detection rules and reward function for the UAV search. Step 3: Establish a mapping area Establish a regional information credibility graph, and propose an optimization problem that combines target search and information collection to maximize the update of the regional credibility graph; The area information trust graph is expressed as: , Wherein, is the regional credibility map corresponding to the grid at time t, and is the grid formed by rasterizing the area . If the grid belongs to the threat area , the regional credibility map is set to 0. is the change in the regional credibility map after the UAV collects information at time t, and is the change in the regional credibility map after the UAV searches for the target at time t. The optimization problem is expressed as: , In the formula, , represent the horizontal and vertical coordinate ranges of the drone , and indicates that the drone flies within the area . Step 4: Design an improved NRO-QMIX algorithm to solve the optimization problem, obtain the optimal solution of the execution actions of each UAV, and complete the multi-UAV collaborative target search and information collection in the urban environment. The improved NRO-QMIX algorithm includes separately processing the observation information of each UAV to generate a local policy network and mixing all UAV locals to generate a global mixing network. By constructing the constraint relationship between the global and the local , a mapping relationship between the global maximum value and the local maximum value is established to ensure that the local optimal action of each UAV is also the global optimal action. Among them, the output layer of the policy network is the NROWAN noise network; The mapping relationship is expressed as: , The constraint relationship to satisfy the mapping relationship is expressed as: , In the formula, local , global , represents the historical actions and observations of the UAV itself, represents the action that the UAV executes at the current moment predicted by the policy network, is the total number of UAVs; The NROWAN noise layer network structure is expressed as: , In the formula, is a learnable parameter, is noise.

2. The multi-UAV collaborative target search method for urban environment based on NRO-QMIX according to claim 1, wherein The multi-UAV search environment architecture described in Step 1 includes an environment model, a target and threat area model, and a UAV model. Among them, the method for constructing the environmental model is as follows: the area to be searched After rasterization, it is divided into multiple squares, respectively representing the number of rows and columns of the squares; the squares are defined as and the center of the square is represented as where, , ; The method for constructing the target and threat area model is as follows: Assume that each target occupies only one square, and there is only one target in each square, and define an event indicating whether there is a target in the square, indicating that there is a target in the current square, indicating that there is no target in the current square. In the unknown area, the preset probability of the existence of a target in each square is 0.5; The threat area is defined as a circle centered at the center position of a certain grid square, and is expressed as: , In the formula, is the center of the threat area, is the threat area number, is the radius of the threat area, R + represents positive real numbers; During the operation of the UAV, it needs to maintain a certain distance from the threat area, and is expressed as: , Wherein, is the coordinate of the drone , and is the preset safety distance between the drone and the dangerous area; The method for constructing the UAV model is: The UAV makes particle motion, and the kinematics of a single UAV is expressed as: , Wherein, is the speed, is the heading angle, is the angular velocity; Single UAVs form a multi-UAV system, which is expressed as: , In the formula, is the total number of drones; Assume that the UAV moves only one grid square at each moment, and the mathematical relationship of the anti-collision positions among multi-UAVs is: , Wherein, is the preset safety distance for anti-collision between multiple unmanned aerial vehicles, , are the coordinates of the unmanned aerial vehicle , j.

3. The multi-UAV collaborative target search method for urban environment based on NRO-QMIX according to claim 1, characterized in that, The design method of the UAV search target detection rules described in Step 2 is: Let event indicate that at time t, the UAV detects a target in the grid . Let indicate that at time t, the UAV does not detect a target in the grid. Then the probability that the UAV at time t detects a target in the grid is expressed as: , In the formula, is the detectivity, is the false alarm rate, is the probability of missed detection, is the true probability of missed detection, represents that there is a target in the grid and represents that there is no target in the grid . Grid The method for updating the target presence probability in it is as follows: , In the formula, is the probability of detecting the presence of a target by the UAV at time t in the grid ; Drone Judge according to the probability value of the target in the current grid to determine the event Value: , When the target existence probability in the grid exceeds the threshold the UAV determines that there is a target in the grid ​ 4. The method for multi-UAV collaborative target search in urban environment based on NRO-QMIX according to claim 1, wherein The reward function described in Step 2 includes a target reward, an information collection reward, an obstacle avoidance penalty, and a time step penalty. Among them, the target reward represents the reward given when the UAV searches for a new target, and the UAV is only rewarded when each target is first discovered. The target reward function is expressed as: , Among them, is the coefficient for adjusting the target reward size, is the total number of drones, is the different number of the drone, is the drone at time t in the grid the probability of detecting the presence of the target; is the threshold of the probability of the target presence; The information collection reward means that within a limited time, the UAV is made to fly over as many grid squares as possible. The information collection reward function is expressed as: , Wherein, is the coefficient for adjusting the size of the information collection reward, is the number of times the grid is searched by the drone at time t; The obstacle avoidance penalty means to avoid collisions among UAVs during the mission. The obstacle avoidance penalty function is expressed as: , Wherein, is the coefficient for adjusting the penalty size of the collision danger zone of the UAV, is the coefficient for adjusting the penalty size of the collision between UAVs; represents the threat zone, s is the threat zone number, 、 are the coordinates of the UAV 、j, is the preset safety distance between the UAV and the danger zone, is the preset safety distance for anti-collision between multiple UAVs; The time step penalty means that the flight time of the UAV is fixed. The time step penalty function is: , wherein, is a coefficient for adjusting the penalty size of the exceeded time step; The reward obtained by all UAVs is the total reward, and the total reward function is expressed as: 。 5. The method for multi-UAV collaborative target search in urban environment based on NRO-QMIX according to claim 1, characterized in that, As described in Step 3 It represents the change in the regional trust map after the UAV information is collected at time t, which is expressed as: , where is the number of times the grid is searched at time t, is the initial regional information credibility of the grid ; The change in the regional credibility map after the UAV searches for the target at time t, which is expressed as: , In the formula, event is that at time t, the UAV detects a target in the grid , means that at time t, the UAV does not detect a target in the grid , is the probability of the existence of a detected target by the UAV at time t in the grid .​ 6. The multi-UAV collaborative target search method for urban environment based on NRO-QMIX according to claim 1, wherein, The data processing process of the improved NRO-QMIX algorithm described in Step 4 is: The unmanned aerial vehicle (UAV) The current observed value And the action at the previous moment Are input into the policy network of each agent network for processing, and the individual action-value function is output , that is, the local ; Among them, Represents the UAV Its own historical actions and observed values, expressed as: , Represents the action that the UAV predicted by the policy network Executes at the current moment; Input the local and global observations s of each drone into the hybrid network to generate the joint action-value function, i.e., the global which is denoted as: and is represented as: which is denoted as: , In the formula, is the hybrid network parameter.

7. The method for multi-UAV collaborative target search in urban environment based on NRO-QMIX according to claim 1, wherein, The process of the improved NRO-QMIX algorithm to solve the optimization problem is: (1)At each discrete moment, the UAV interacts with the environment of the exploration area. At time t, the UAV generates its own observation value at the current moment , and inputs the own observation value and the action at the previous moment as well as the global observation into the NRO-QMIX neural network, and outputs the action with the maximum value of the global and local for each UAV . The UAV executes this action and obtains the total reward of multiple UAVs at the current moment t , thus entering the next moment and completing a complete interaction loop;​ (2) The drones start a new cycle, continue to interact with the environment, and continuously accumulate rewards. After a finite number of time steps, all cycles end, and the overall multi-drone , during several rounds of training, the NRO-QMIX neural network is continuously updated according to the previous training experience and the designed loss function until the training is completed; (3) Use the trained NRO-QMIX neural network to obtain the optimal action performed by the UAV under the current observation.

8. The method for multi-UAV collaborative target search in urban environment based on NRO-QMIX according to claim 7, wherein The loss function is expressed as: , In the formula, are the parameters of the target network, reflects the noise level, is the proportionality coefficient, is the discount coefficient, is the total reward of multiple UAVs at time t, is the target global noise value function, is the global noise value function; Among them, , , In the formula, is the input dimension of the last layer of the NROWAN noise network, is the number of output actions, is the final expected value of, is the reward at the current time t, is the maximum reward, is the minimum reward, is the noise variance of the weight, is the noise variance of the bias.

Citation Information

Patent Citations

  • Unmanned aerial vehicle cluster multi-target search method and system based on deep reinforcement learning

    CN112947575A

  • Unmanned aerial vehicle cluster collaborative search method and system based on deep multi-agent reinforcement learning

    CN119292342A