Unmanned cluster distributed collaborative decision-making method for complex search scene

By introducing three-dimensional spatiotemporal grid modeling, multi-view target consistency judgment and reinforcement learning decisions into the drone cluster system, path allocation and coordinated scheduling are optimized, and path planning and coordinated decision-making in complex search tasks are solved, and efficient and flexible drone collaborative searches are achieved.

CN120447576APending Publication Date: 2025-08-08NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510569024.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing UAV cluster collaborative system lacks the ability to model dynamic target behavior in complex search tasks, and cannot effectively integrate the uncertainty of target movement, which makes it difficult to optimize the path allocation globally and lacks the utilization of multi-view information, which is prone to repeated identification or misidentification, and the roundup strategy lacks real-time and elasticity, making it difficult to adapt to target movement or dynamic adjustment of strategy, and non-consensus drones lack efficient collaborative scheduling mechanisms, which are prone to cause task conflicts or resource waste.

Method used

By modeling the search area into a three-dimensional spatiotemporal grid, combining the target state transition probability and the UAV energy consumption model, a comprehensive hit rate model is established, and a simulated annealing algorithm is used to optimize path allocation; based on multi-view overlap and motion consistency judgment, a reinforcement learning decision model is designed; a consensus state collaborative prediction and dynamic graph scheduling mechanism is constructed to optimize round-up point allocation and conflict resolution.

Benefits of technology

It significantly improves the success rate of search in complex environments, reduces misidentification and repeated identification, realizes autonomous decision-making and coordinated scheduling of drones, adapts to dynamic adjustment needs, reduces resource waste, and improves search efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120447576A_ABST
    Figure CN120447576A_ABST
Patent Text Reader

Abstract

The invention discloses a complex search scene-oriented unmanned cluster distributed collaborative decision-making method, which comprises the following steps of: modeling a search area into a three-dimensional space-time grid, establishing an unmanned aerial vehicle cluster comprehensive hit rate model and a cost function, and solving approximate optimal path distribution; establishing a target consistency judgment model based on multi-view space overlapping, and judging whether observation targets are the same target or not; target encircle collaborative decision-making is realized by improving reinforcement learning; forming a consensus unmanned aerial vehicle cluster based on group consensus, and performing position and speed estimation on a target in combination with an unmanned cluster observation error and target steering; cooperative network graph relation construction is carried out on non-consensus unmanned aerial vehicles through a dynamic graph theory, so that space-time scheduling is carried out, and surrounding point distribution and conflict resolution are optimized. According to the method, through path optimization, target consistency judgment, reinforcement learning of a hunting strategy, consensus target state collaborative prediction and a dynamic graph scheduling mechanism, efficient collaborative hunting and task allocation of multiple unmanned aerial vehicles in a complex environment are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an unmanned cluster distributed collaborative decision-making method for complex search and capture scenarios, and belongs to the field of unmanned aerial vehicle collaborative search decision-making and allocation scheduling. Background Art

[0002] Current drone swarm collaborative systems still face multiple technical bottlenecks when handling complex search and capture missions. First, most existing path planning methods lack the ability to model dynamic target behavior and are unable to effectively incorporate the uncertainty of target motion. This makes it difficult to achieve a globally optimal path allocation, impacting capture efficiency. Second, when multiple drones observe the same target, existing methods are relatively crude in terms of perspective fusion and target consistency judgment. They lack mechanisms to fully utilize multi-perspective information for judgment, which can easily lead to duplicate or misidentification issues. Furthermore, existing capture strategies often rely on preset rules or centralized control, lacking flexibility and real-time performance, making them difficult to adapt to target movement or dynamic strategy adjustments. Furthermore, after the swarm reaches consensus, existing systems lack a collaborative prediction mechanism for target state (such as position and velocity), making it difficult to integrate asynchronous observation information from multiple drones. This leads to deviations in prediction accuracy, affecting the scheduling efficiency and capture effectiveness of subsequent non-consensus drones. Furthermore, after receiving information, non-consensus drones lack an efficient collaborative scheduling mechanism for rational division of labor, which can easily lead to task conflicts and waste resources. Therefore, how to build a full-process method system covering path optimization, target consistency judgment, collaborative capture reinforcement learning decision-making, consensus state prediction and non-consensus scheduling optimization is a key issue that urgently needs to be broken through in the current drone swarms in complex search and capture tasks. Summary of the Invention

[0003] Purpose of the invention: In order to overcome the shortcomings of the existing technology, the present invention provides an unmanned swarm distributed collaborative decision-making method for complex search and capture scenarios. By introducing spatiotemporal grid modeling, path and energy consumption optimization, multi-perspective target consistency judgment, consensus state collaborative prediction and dynamic graph scheduling mechanism, it can achieve collaborative decision-making of multiple UAVs and efficient capture of targets in a dynamic environment.

[0004] Technical solution: To achieve the above purpose, the technical solution adopted by the present invention is:

[0005] A distributed collaborative decision-making method for unmanned swarms in complex manhunt scenarios includes the following steps:

[0006] Step 1: Model the search area as a three-dimensional space-time grid, define the drone cluster attributes and target state transition probability, and then establish a drone energy consumption model and a drone cluster comprehensive hit rate model. Combined with the flight path, time and energy consumption constraints, a total cost function is constructed, and the objective function is minimized through the simulated annealing algorithm to solve the approximate optimal path allocation.

[0007] Step 2: Based on the obtained approximate optimal path allocation, the UAV perspective overlap rate and motion consistency constraints are used to establish a target consistency judgment model based on multi-perspective spatial overlap to determine whether the observed targets are the same target.

[0008] Step 3: Design the state space, action space, and reward function based on the target consistency judgment model, and realize the target capture collaborative decision-making through improved reinforcement learning.

[0009] Step 4: Based on the target capture collaborative decision-making, a consensus drone cluster is formed based on group consensus, and the position and speed of the target are estimated by combining the observation error of the drone cluster and the target steering.

[0010] Step 5: Based on the consensus UAV cluster’s estimation of the target’s position and velocity, a collaborative network graph relationship is constructed for the non-consensus UAVs through dynamic graph theory to perform spatiotemporal scheduling, optimize the allocation of capture points, and resolve conflicts.

[0011] Preferred: The target consistency judgment model is:

[0012]

[0013] Among them, τ vol , τ area is the threshold of volume overlap ratio and projection overlap ratio. This constraint indicates that the view overlap must reach a certain level before the identified drones are considered to be the same. is the distance difference between UAV j1 and UAV j2. The integral term reflects the target in the time window The maximum possible displacement within. The maximum flight speed of the UAV. The upper and lower limits of the time integration range

[0014] Optimal: The comprehensive hit rate model of the drone cluster is:

[0015]

[0016] in, is the comprehensive hit rate.

[0017] Preferably, step 4 includes the following steps:

[0018] Step 41: When the drone cluster recognizes the target at the same time, in order to reach a consensus, the consensus conditions of the drones that recognize the target are determined. First, the observation state of each drone i is expressed as follows:

[0019]

[0020] Among them, the states are the position vector, velocity vector, and angle of drone j in sequence.

[0021] If there are N identify UAVs recognize the target, and each UAV has its own recognition confidence p j , then for the cluster, the target is first judged by the consistency observation model, and then the spatiotemporal joint perspective overlap confidence ζ is obtained TS :

[0022]

[0023] Among them, N identify is the number of drones that have identified the target in the cluster. p1,…, μ is the recognition probability of each drone in the drone cluster. TS is the time decay coefficient, which controls the weight of historical data. The time difference between any two UAVs j1 and j2 in identifying the target. By introducing the time decay term Reduce the confidence of asynchronous observations to avoid interference from stale data.

[0024] The group reaches consensus if the spatiotemporal joint overlap confidence satisfies the following conditions:

[0025] ζ TS >ζ th

[0026] Among them, th is the consensus threshold.

[0027] After forming a consensus cluster, the observation angle of each drone is set to a ray pointing from the drone position to the target, and the equation is:

[0028]

[0029] Step 42: Optimize the target position estimate using weighted least squares method.

[0030] First, define the observation residual of UAV j That is, the vertical distance from the actual position of the target to the jth observation ray, which is used to measure the error of the observation data:

[0031]

[0032] It can be seen that the smaller the residual, the closer the actual position of the target is to the observation ray of UAV j. So we can get:

[0033]

[0034] in, is the observation weight, which is related to the observation accuracy.

[0035] To solve this equation, we need to take the partial derivative and set the derivative to zero to obtain a system of linear equations:

[0036]

[0037] Arranged into matrix form, we can get:

[0038]

[0039] Among them are:

[0040]

[0041] Therefore, the predicted position of the target can be obtained, and the solution is:

[0042]

[0043] Step 43, calculate the target's velocity: Assuming that both the drone and the target are moving at a constant speed, the target's apparent velocity vector relative to drone i is for:

[0044]

[0045] in, is the target k velocity vector.

[0046] Step 44: The UAV continuously observes the target. When the target rotates, the change in azimuth angle is determined by the velocity component v perpendicular to the line of sight. ⊥ Caused by the unit vector n perpendicular to the line of sight j for:

[0047]

[0048] Therefore, when UAV j observes the target, the target's azimuth angle change rate is for:

[0049]

[0050] in, is the distance between UAV j and target k. The azimuth angle change rate quantifies the impact of target motion on the observation angle and provides a basis for velocity estimation.

[0051] Step 45, As unknowns, construct a system of linear equations:

[0052]

[0053] Similarly, writing it in matrix form:

[0054]

[0055] Among them are:

[0056]

[0057] Since the above matrix C observe It is not a square matrix, so the least squares solution is:

[0058]

[0059] Among them, W observe is a diagonal weight matrix.

[0060] In step 46, after the drone cluster reaches consensus, it completes the collaborative prediction of the target through group consensus. The position vector and velocity vector of the target are as follows:

[0061]

[0062] in, represents the position vector of the target, Represents the velocity vector.

[0063] An unmanned swarm distributed collaborative decision-making system for complex manhunt scenarios, used to implement the unmanned swarm distributed collaborative decision-making method for complex manhunt scenarios, includes an optimal path unit, a target consistency judgment model unit, an improved reinforcement learning unit, a target position and speed estimation unit, and a collaborative network unit, wherein:

[0064] The optimal path unit is used to model the search area as a three-dimensional space-time grid, while defining the drone cluster attributes and target state transition probabilities, and then establish a drone energy consumption model and a drone cluster comprehensive hit rate model. The total cost function is constructed by combining the flight path, time and energy consumption constraints, and the objective function is minimized through the simulated annealing algorithm to solve the approximate optimal path allocation.

[0065] The target consistency judgment model unit is used to establish a target consistency judgment model based on multi-perspective spatial overlap according to the obtained approximate optimal path allocation through the drone perspective overlap rate and motion consistency constraint, and determine whether the observed targets are the same target.

[0066] The improved reinforcement learning unit is used to design the state space, action space and reward function based on the target consistency judgment model, and realize the target capture collaborative decision-making through improved reinforcement learning.

[0067] The target position and speed estimation unit is used to form a consensus drone cluster based on group consensus according to the target capture collaborative decision, and estimate the position and speed of the target in combination with the unmanned cluster observation error and target steering.

[0068] The collaborative network unit is used to estimate the position and speed of the target based on the consensus drone cluster, and to construct a collaborative network graph relationship for non-consensus drones through dynamic graph theory, so as to perform spatiotemporal scheduling, optimize the allocation of capture points and resolve conflicts.

[0069] Compared with the prior art, the present invention has the following beneficial effects:

[0070] 1. This method models the search area as a three-dimensional space-time grid. Combining target state transition probabilities with a drone energy consumption model, it constructs a total cost function (integrating path length, time, and energy consumption) and uses a simulated annealing algorithm to find a near-optimal path allocation. Coverage constraints and communication synchronization constraints ensure a complete global search, significantly reducing resource waste and improving the success rate of target searches in unknown environments.

[0071] 2. This invention establishes a multi-view target consistency judgment model by defining volume overlap ratios, projection overlap ratios, and motion consistency constraints, reducing the problems of misidentification and duplicate identification. Based on reinforcement learning, the state space, action space, and reward function are designed to guide drones in autonomously determining capture strategies. Furthermore, consensus clusters are used to collaboratively predict target position and velocity, and dynamic graph theory is used to construct a collaborative network of non-consensus drones. Combined with communication quality constraints, spatiotemporal coordinated capture optimization is achieved. This layered collaborative mechanism balances real-time performance with flexibility, adapting to the dynamic adjustment requirements in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 This is a flowchart of the unmanned cluster collaborative search planning of the present invention.

[0073] Figure 2 This is a flowchart of the unmanned cluster autonomous decision-making based on reinforcement learning in the present invention.

[0074] Figure 3 This is a flowchart of the unmanned cluster collaborative capture based on dynamic graph theory of the present invention.

[0075] Figure 4 This is the unmanned cluster collaborative search planning experiment diagram of the present invention.

[0076] Figure 5 This is a diagram of the non-consensus UAV spatiotemporal dynamic capture and scheduling experiment of the present invention. DETAILED DESCRIPTION

[0077] The present invention is further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0078] This embodiment proposes a distributed collaborative decision-making method for unmanned swarms for complex search and capture scenarios. The search area is modeled as a three-dimensional space-time grid, and the drone cluster attributes and target state transition probabilities are defined. Furthermore, a drone energy consumption model and a drone cluster comprehensive hit rate model are established, and a total cost function is constructed by combining flight path, time and energy consumption constraints. The objective function is minimized through a simulated annealing algorithm to solve the approximate optimal path allocation. Through the drone perspective overlap rate and motion consistency constraints, a target consistency judgment model based on multi-perspective spatial overlap is established to determine whether the observed targets are the same target. The state space, action space and reward function are designed, and target capture collaborative decision-making is achieved through a new type of reinforcement learning. A consensus drone cluster is formed based on group consensus, and the position and velocity of the target are estimated by combining the unmanned cluster observation error and target steering. Based on the consensus drone cluster's estimate of the target's position and velocity, a collaborative network graph relationship is constructed for non-consensus drones through dynamic graph theory, thereby performing spatiotemporal scheduling, optimizing capture point allocation and conflict resolution.

[0079] like Figure 1 、 2 , 3, the specific steps are as follows:

[0080] Step 1: Model the search area as a three-dimensional space-time grid, defining the drone swarm attributes and target state transition probabilities. Furthermore, establish a drone energy consumption model and a comprehensive hit rate model for the drone swarm. Combined with flight path, time, and energy constraints, a total cost function is constructed. A simulated annealing algorithm is used to minimize the objective function and obtain a near-optimal path assignment.

[0081] First, the forest search area is modeled as a three-dimensional space-time grid G:

[0082] G={(x,y,z,t)|x∈[1,M G ],y∈[1,N G ],z∈[1,H G ],t∈[1,T G ]}

[0083] Where (x, y, z, t) are the x-coordinate, y-coordinate, z-coordinate and time step of the search area respectively. G 、N G 、H G They are the maximum values in the x-coordinate direction, y-coordinate direction, and z-coordinate direction of the region, and the total number of two-dimensional grids is K G =M G ×N G . T G The maximum time of the region.

[0084] For a drone swarm performing a search and capture mission, it is represented as follows:

[0085]

[0086] Among them, N arrest is the number of drones in the drone swarm.

[0087] And for each drone Contains the following properties:

[0088]

[0089] in, is the flight speed of the drone. The flight direction of the drone. is the current position of the drone. The flight time of the drone. is the detection radius of the UAV.

[0090] For the target in G, its specific position cannot be known. Set the target k to the speed and direction Move, then the state transition of the target can be modeled through the Markov process, and the prior probability distribution of drone j finding target k at node (x, y) can be obtained:

[0091]

[0092] Among them, P j (x,yt, search ) represents the UAV j at time t sear The probability that node (x, y) discovers target k. Table ) shows the UAV j at time t search -1 The probability that node (x′, y′) discovers target k. It represents the variance of the target position change and reflects the motion uncertainty of the target.

[0093] At the same time, since the energy consumption limit of the UAV may affect its flight distance, a UAV energy consumption model is established:

[0094]

[0095] in, is the UAV flight energy consumption, E sense Sensing energy consumption for drones. And there are:

[0096]

[0097] in, is the current speed of the UAV. sensor is the sensor energy consumption, which is a fixed value. is the energy consumption coefficient, which is used to indicate the impact of this factor on energy consumption.

[0098] After establishing the UAV energy consumption model, we can further define the total cost of UAV cluster collaborative search planning

[0099]

[0100] in, is the flight path length of the j-th UAV. is the flight time of the jth UAV. is the flight energy consumption of the j-th UAV search process. is the weight coefficient.

[0101] Taking into account the target's movement and time decay, the comprehensive hit rate Defined as the probability that a cluster sees a target at least once:

[0102]

[0103] The structure of this formula is a multi-product form, which represents the probability calculation of multiple attempts. The inner product is the probability of drone i at a given position (x, y) and time t search Under this condition, the probability of not finding the target is multiplied, The outer product represents the multiplication of the probability of the target at all positions being undetected at each time t. The final total hit rate is is from all moments t search The product of the probability of finding the target at least once is obtained by multiplying the probability of finding the target at all points (x, y). Since it is in product form, it means that in order to successfully find the target at least once, the cluster must search and try at all time steps and all positions, and then calculate the probability of finding the target at least once by inverting the overall probability.

[0104] Furthermore, in order to jointly minimize the total cost and maximize the comprehensive hit rate, it is necessary to minimize the objective function:

[0105]

[0106] in, Used to balance target dimensions.

[0107] At the same time, the constraints are also expressed as follows:

[0108] 1. Battery life constraints

[0109] The flight time of each drone cannot exceed the endurance time:

[0110]

[0111] 2. Path continuity

[0112] The flight path of each drone must be continuous to avoid jumping:

[0113]

[0114] in, The UAV at time t search 's coordinates. The UAV at time t search +1 coordinate. Δt search is the time step.

[0115] 3. Coverage constraints

[0116] All grids are covered by at least one drone:

[0117]

[0118] in, is the grid area in G, is the coverage area of UAV j. is an indicator function. When the condition is met, the value of the indicator function is 1, otherwise it is 0. This formula is used to ensure that the entire search area is covered.

[0119] 4. Communication synchronization constraints

[0120] The drone must maintain a connection with at least one other drone within the communication range to ensure data synchronization:

[0121]

[0122] in, are the positions of UAVs j1 and j2 respectively. comm is the shortest communication distance between drones.

[0123] After completing the construction of the optimization problem, considering the strong global search properties of the optimization problem and the fact that it is a multi-objective optimization problem, the simulated annealing algorithm is used to search for an approximate optimal solution:

[0124] First, initialize: generate an initial UAV cluster search path plan to ensure that the path meets the above path constraints, path continuity, coverage constraints and communication synchronization constraints. Set a higher initial temperature Allow the algorithm to accept poor solutions in the early stages. Develop a cooling plan to determine the strategy for temperature reduction:

[0125] T next =α sa ·T current ,αsa <1

[0126] Among them, T current 、T next are the current and next temperatures respectively. sa is the coefficient for controlling the drop.

[0127] Next, calculate the objective function value of the initial solution, the objective function f sa The definition is as follows:

[0128]

[0129] in, is the weight coefficient, which is used to balance the total cost and the comprehensive hit rate.

[0130] Furthermore, iterative optimization begins. In each round of iteration, the simulated annealing algorithm searches through the following steps:

[0131] 1. Generate a new solution: Based on the current solution, a new solution is generated through random perturbations. The perturbation method is to randomly adjust the flight path of a certain drone to ensure that the new solution still meets all constraints.

[0132] 2. Calculate the objective function value: Calculate the objective function value based on the new solution.

[0133] 3. Accept the new solution: Decide whether to accept the new solution based on the Metropolis criterion: If the objective function value of the new solution is better than the current solution, then accept the new solution. If the objective function value of the new solution is worse, then accept the new solution based on the probability Accept the new solution, where Δf sa is the difference between the objective function value of the new solution and the current solution.

[0134] 4. Update temperature: Lower the temperature according to the cooling plan, gradually reducing the probability of accepting a worse solution.

[0135] When the temperature drops to a certain threshold Or when the maximum number of iterations is reached, the algorithm terminates and outputs the current optimal solution, which is the optimal search path solution for the drone cluster.

[0136] Step 2: Based on the UAV perspective overlap rate and motion consistency constraints, a target consistency judgment model based on multi-perspective spatial overlap is established to determine whether the observed targets are the same target.

[0137] In a complex search and rescue scenario, when multiple drones identify a target at the same time, the sensor perspective needs to be modeled first. The perspective coverage area of each drone j is Defined as a 3D cone:

[0138]

[0139] in, Detect distance for drones.

[0140] Next, we calculate the perspective overlap rate and define the perspective overlap area between UAV j1 and UAV j2. is the intersection of the two 3D cones:

[0141]

[0142] Furthermore, the overlap metric volume overlap ratio can be obtained Projection overlap ratio

[0143]

[0144] Where Vol(·) is the volume, which indicates the consistency of the overlap of the three-dimensional coverage area. Area(Proj(·)) is the area of the two-dimensional plane projection of the region, which indicates the consistency of the two-dimensional plane coverage overlap.

[0145] At the same time, if there is an observation time difference Δt between the two drones op :

[0146]

[0147] in, are the times when UAV j1 and UAV j2 observe the target respectively.

[0148] The volume overlap ratio after attenuation is Projection overlap ratio for:

[0149]

[0150] Among them, μ op is the attenuation coefficient.

[0151] Therefore, based on the above target consistency judgment rule, if UAV j1 and UAV j2 identify a target and meet the following conditions, they are determined to be the same target:

[0152] 1. Spatial overlap constraints

[0153]

[0154] Among them, τ vol , τ area is the threshold of volume overlap ratio and projection overlap ratio. This constraint indicates that the view overlap must reach a certain level before the identified drones are considered to be the same.

[0155] 2. Motion consistency constraints

[0156]

[0157] in, is the distance difference between UAV j1 and UAV j2. The integral term reflects the target in the time window The maximum possible displacement within. The maximum flight speed of the UAV. The upper and lower limits of the time integration range The calculation is as follows:

[0158]

[0159]

[0160] This constraint means that the target must always be in the view of all drones. For multiple drones, the above model can be used to determine target consistency between each pair.

[0161] Step 3: Design the state space, action space, and reward function, and implement collaborative decision-making for target capture through novel reinforcement learning.

[0162] Based on the above judgment model, we start to establish a collaborative hunting method based on multi-agent deep reinforcement learning. First, assume that UAV j1 detects target k during the search process, and after obtaining relevant detection information, it maintains the same speed as the target. During the flight, continuous group broadcasting is carried out. The broadcast information package is as follows

[0163]

[0164] in, The ID number of the drone j1. The time when UAV j1 detects the target. is the position of UAV j1 when it detects the target. j1 is the recognition probability of the target detected by UAV j1. The moving direction of UAV j1 when it detects the target. is the real-time speed of UAV j1. The number of times UAV j1 has been assisted.

[0165] For the idle drone j2, it receives the information packet broadcast by drone j1 After that, the information in the information packet is processed to obtain

[0166]

[0167] in, is the spatiotemporal alignment distance between UAV j1 and UAV j2. is the moving speed of target k. is the moving direction of target k. is the location of target k. is the recognition probability of the target by UAV j1 after processing.

[0168] During the overall search process, the drone continuously tracks a target after discovering it. Therefore, we assume that the drone's flight speed and direction are consistent with the target's movement. This assumption enables the drone to accurately predict the target's path. However, considering packet transmission delays, we need to predict the target's real-time position and state to adjust the recognition confidence. As the drone continues to fly, its speed and packet transmission delays will affect recognition confidence.

[0169] Specifically, the faster the drone flies, the worse the recognition results may be. This is because the drone faces greater dynamic changes during the tracking process, resulting in inaccurate real-time recognition information. Furthermore, packet transmission delays can cause the information received by the drone to lag behind the target's actual location and status, increasing recognition errors. When packets are transmitted from a distant drone, the longer the delay, the lower the confidence level of the recognition results. To compensate for these effects, the target's current location must be predicted and the recognition confidence level adjusted based on the transmission delay to improve overall search accuracy. Therefore, the following are the key factors:

[0170]

[0171] in, is the information packet transmission delay between UAV j1 and UAV j2. safe The safe tracking distance of the drone. fusion , β fusion is the adjustment factor for speed and distance, which determines the degree of their influence on the recognition probability. And it can be calculated as follows:

[0172]

[0173] in, The time when drone j2 receives the information packet.

[0174] After the above preprocessing, the information packet is obtained. After that, the reinforcement learning is improved. First, the state space S is designed, where the state vector s t for:

[0175]

[0176] Furthermore, the action vector a in the action space A t for:

[0177]

[0178] Among them, pos desired is the expected capture position. desired is the output speed. desired is the output flight angle. desired is the expected arrival time. is the assistance intensity for UAV j1, representing the willingness to assist.

[0179] Furthermore, the reward function reward is designed as:

[0180]

[0181] in, The value is 1 if the capture is successful, otherwise it is 0. pred To predict the capture location, t pred is the predicted arrival time. r , β r , γ r , δ r are weight coefficients respectively. For position matching and time matching, higher assistance intensity leads to stricter matching accuracy requirements. Regarding assistance intensity rewards, moderate increases in assistance intensity are encouraged to avoid excessive dominance by a single drone.

[0182] For each predictor variable in the reward function, the following requirements must be met:

[0183]

[0184]

[0185] Among them, N request The number of drones making the request.

[0186] Furthermore, a list of requested drones is formed according to the strength of assistance. assist The drones in the list are arranged in descending order of assistance strength:

[0187] List assist ={j|φ j >φ j+1},j=1,2…,N request

[0188] Therefore, UAV j2 has the following choices:

[0189]

[0190] Among them, N request is the number of drones that sent information packets. N assist The number of assisting drones required for a drone to capture a target. assist represents assisting the current drone j. search represents a search operation. The above decision-making process indicates that when the highest assist intensity output by each idle drone through reinforcement learning corresponds to a number of drones currently assisting less than the number required for capture, the idle drone decides to assist. Otherwise, a recursive decision is made until either assisting or searching is completed.

[0191] Furthermore, the drone that decides to assist in the roundup responds to the requesting drone, and the requesting drone increases the number of assistance after receiving the response:

[0192]

[0193] in, The current number of assisted drones that have been assisted. After the required number of roundups is met, the drone is requested to stop broadcasting information packets.

[0194] Furthermore, when multiple drones identify a target, they first use the consistency judgment model to determine whether they have identified the same target, and then group the drones that have identified the same target:

[0195]

[0196] Among them, group m ,m∈[1,…,N goal ] is the UAV group m that recognizes the same target. N goal is the target number.

[0197] The number of assists for each drone in the group remains the same at all times, namely:

[0198]

[0199] At the same time, the initial assistance quantity of each drone is:

[0200]

[0201] Among them, num(group m ) is group m The number of drones.

[0202] Next, each drone in the group broadcasts an information packet. After the other idle drones receive it, since each idle drone receives information packets from multiple drones, they calculate the strength of each drone sending an information packet and compare them according to the above decision. They select the one with the greatest strength to assist in the roundup and send a response at the same time.

[0203] Step 4: Based on group consensus, a consensus drone cluster is formed, and the position and speed of the target are estimated by combining the observation error and target steering of the drone cluster.

[0204] When the drone cluster recognizes the target at the same time, in order to reach a consensus, the consensus conditions of the drones that recognize the target are determined. First, the observation state of each drone i is expressed as follows:

[0205]

[0206] Among them, the states are the position vector, velocity vector, and angle of drone j in sequence.

[0207] If there are N identify UAVs recognize the target, and each UAV has its own recognition confidence p j , then for the cluster, the target is first judged by the consistency observation model, and then the spatiotemporal joint perspective overlap confidence ζ is obtained TS :

[0208]

[0209] Among them, N identify is the number of drones that have identified the target in the cluster. p1,…, μ is the recognition probability of each drone in the drone cluster. TS is the time decay coefficient, which controls the weight of historical data. The time difference between any two UAVs j1 and j2 in identifying the target. By introducing the time decay term Reduce the confidence of asynchronous observations to avoid interference from stale data.

[0210] The group reaches consensus if the spatiotemporal joint overlap confidence satisfies the following conditions:

[0211] ζ TS >ζ th

[0212] Among them, th is the consensus threshold.

[0213] After forming a consensus cluster, the observation angle of each drone is set to a ray pointing from the drone position to the target, and the equation is:

[0214]

[0215] Due to measurement errors, multiple rays may not intersect at the same point, so we use weighted least squares to optimize the target position estimate.

[0216] First, define the observation residual of UAV j That is, the vertical distance from the actual position of the target to the jth observation ray, which is used to measure the error of the observation data:

[0217]

[0218] It can be seen that the smaller the residual, the closer the actual position of the target is to the observation ray of UAV j. So we can get:

[0219]

[0220] in, is the observation weight, which is related to the observation accuracy.

[0221] Furthermore, to solve the equation, we need to take the partial derivative and set the derivative to zero to obtain a system of linear equations:

[0222]

[0223] Arranged into matrix form, we can get:

[0224]

[0225] Among them are:

[0226]

[0227] Therefore, the predicted position of the target can be obtained, and the solution is:

[0228]

[0229] Furthermore, in addition to the target's position coordinates, we also need to calculate the target's velocity. Assuming that both the drone and the target are moving at a constant speed, the target's apparent velocity vector relative to drone i is for:

[0230]

[0231] in, is the target k velocity vector.

[0232] The UAV continuously observes the target. When the target rotates, the change in azimuth is mainly caused by the velocity component v perpendicular to the line of sight. ⊥ Caused by the unit vector n perpendicular to the line of sight j for:

[0233]

[0234] Therefore, when UAV j observes the target, the target's azimuth angle change rate is for:

[0235]

[0236] in, is the distance between UAV j and target k. The azimuth angle change rate quantifies the impact of target motion on the observation angle and provides a basis for velocity estimation.

[0237] Further, As unknowns, construct a system of linear equations:

[0238]

[0239] Similarly, writing it in matrix form:

[0240]

[0241] Among them are:

[0242]

[0243] Since the above matrix C observe It is not a square matrix, so the least squares solution is:

[0244]

[0245] Among them, W observe is a diagonal weight matrix.

[0246] In summary, after the drone cluster reaches consensus, it can complete the collaborative prediction of the target through group consensus, and the target position vector and velocity vector as follows:

[0247]

[0248] Step 5: Based on the consensus UAV cluster’s estimation of the target’s position and velocity, a collaborative network graph relationship is constructed for the non-consensus UAVs through dynamic graph theory to perform spatiotemporal scheduling, optimize the allocation of capture points, and resolve conflicts.

[0249] The consensus drone cluster determines the target's position and speed through group consensus, and sends the target's position and speed to the remaining non-consensus drones. Then, based on dynamic graph theory modeling of the spatiotemporal relationship of the drone cluster, the non-consensus drones are able to collaboratively capture the target through time difference compensation and spatial path optimization.

[0250] First, construct the graph G capture :

[0251] G capture ={V capture ,ε capture}

[0252] Among them, V capture is a node collection, representing all drone nodes. capture is an edge set, representing the connection between nodes. The attributes of drone nodes are as follows:

[0253]

[0254] in, The consensus of each drone node in the drone cluster, including its own position own velocity vector and the time the message was sent Each drone node in the non-consensus drone cluster, including its own location own velocity vector and the time the information was received

[0255] If UAV j1 and UAV j2 can communicate, then there is a communication edge, and the edge weight is the communication quality

[0256]

[0257] Among them, α comm is the communication quality quantization coefficient, is the distance between UAV j1 and UAV j2.

[0258] At the same time, since there is a time difference when the non-consensus drone k receives the consensus information, the time difference needs to be compensated to predict the current position of the target:

[0259]

[0260] in, Predict the location of the target. are the components of the target velocity along the x-axis and y-axis.

[0261] Furthermore, to generate candidate capture points, we first determine the safe capture area with the predicted target location as the center and a radius of r. safe Generate a uniformly distributed set of candidate points P capture :

[0262]

[0263] Next, according to the target's movement direction, the priority of the candidate capture points is adjusted, and the drone is preferentially assigned to the candidate points ahead of the target's movement direction to improve the interception success rate:

[0264]

[0265] in, Candidate capture points. is the weight of the candidate capture point. is the candidate capture point (x capture ,y capture ) relative to the predicted target position The direction angle and the target moving direction angle are:

[0266]

[0267] Next, we build a virtual collaborative network. First, we transform the candidate points into Join the space-time graph G capture , add a virtual edge between each non-consensus drone j2 and the candidate capture point, and the edge weight is the spatiotemporal collaboration cost

[0268] in, are the current position and candidate capture point of UAV k respectively. capture , β capture , γ capture , δ capture is the weight coefficient, satisfying α capture +β capture +γ capture +δ capture = 1. The first path time term in the space-time coordination cost Arrival of drone j2 at the capture point The time cost is 100%, which encourages the selection of UAVs with short distance and fast speed. The second time lag term Penalize drones that receive information late and prioritize drones with updated data. This indicates that drones with speeds close to the target are easier to intercept. The fourth item is the candidate weight item. This indicates that the collaborative cost of high-weight candidate points is reduced, and thus they are preferred.

[0269] Each drone selects the local optimal capture point

[0270]

[0271] in, is the heuristic function, and we have:

[0272]

[0273] Among them, λ crowded is the congestion penalty coefficient. Indicates the congestion of the capture point and counts the currently selected capture point The number of drones should be reduced to avoid drone gathering as much as possible.

[0274] After pre-allocation is completed, if multiple drones still choose the same capture point The auction mechanism is triggered. The conflicting drones broadcast their bids. The drone with the highest bid wins the roundup spot

[0275]

[0276] The formula shows that the faster the drone is and the closer it is to the capture point, the higher the bid. At the same time, the weight of the candidate point is taken into account. When the conflict is resolved, the bid of the high-weight candidate point is higher. Increase, so that high-weight candidate points are more easily assigned, ensuring the optimal interception direction.

[0277] The drones that fail the auction immediately fall back to the suboptimal candidate point according to the heuristic function sorting and recalculate the spatiotemporal coordination cost. If the suboptimal point is still occupied, a local broadcast is triggered to notify other drones to adjust their paths to avoid cascading conflicts. At the same time, considering that distance affects communication quality and makes it impossible to send information accurately, constraints are also required when allocating points:

[0278]

[0279] in, To ensure basic communication quality. capture is the number of drones participating in the roundup. It is necessary to ensure that when drone j is assigned, there is at least one consensus drone in the vicinity that maintains communication with it.

[0280] In another embodiment of the present invention, an unmanned swarm distributed collaborative decision-making system for complex search and arrest scenarios is provided, which is used to implement the unmanned swarm distributed collaborative decision-making method for complex search and arrest scenarios, including an optimal path unit, a target consistency judgment model unit, an improved reinforcement learning unit, a target position and speed estimation unit, and a collaborative network unit, wherein:

[0281] The optimal path unit is used to model the search area as a three-dimensional space-time grid, while defining the drone cluster attributes and target state transition probabilities, and then establish a drone energy consumption model and a drone cluster comprehensive hit rate model. The total cost function is constructed by combining the flight path, time and energy consumption constraints, and the objective function is minimized through the simulated annealing algorithm to solve the approximate optimal path allocation.

[0282] The target consistency judgment model unit is used to establish a target consistency judgment model based on multi-perspective spatial overlap according to the obtained approximate optimal path allocation through the drone perspective overlap rate and motion consistency constraint, and determine whether the observed targets are the same target.

[0283] The improved reinforcement learning unit is used to design the state space, action space and reward function based on the target consistency judgment model, and realize the target capture collaborative decision-making through improved reinforcement learning.

[0284] The target position and speed estimation unit is used to form a consensus drone cluster based on group consensus according to the target capture collaborative decision, and estimate the position and speed of the target in combination with the unmanned cluster observation error and target steering.

[0285] The collaborative network unit is used to estimate the position and speed of the target based on the consensus drone cluster, and to construct a collaborative network graph relationship for non-consensus drones through dynamic graph theory, so as to perform spatiotemporal scheduling, optimize the allocation of capture points and resolve conflicts.

[0286] The unmanned cluster collaborative search planning experiment of the present invention is as follows Figure 4 As shown in the figure, the non-consensus UAV spatiotemporal dynamic capture scheduling experiment of the present invention is as follows Figure 5 The present invention achieves efficient collaborative capture and task allocation of multiple UAVs in complex environments through path optimization, target consistency judgment, reinforcement learning capture strategy, consensus target state collaborative prediction and dynamic graph scheduling mechanism.

[0287] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A distributed collaborative decision-making method for unmanned swarms in complex search and arrest scenarios, characterized by: The steps include: Step 1: Model the search area as a three-dimensional space-time grid, define the drone cluster attributes and target state transition probabilities, and then establish a drone energy consumption model and a drone cluster comprehensive hit rate model. Combined with the flight path, time, and energy consumption constraints, a total cost function is constructed. The objective function is minimized using a simulated annealing algorithm to solve the approximate optimal path allocation. Step 2: Based on the obtained approximate optimal path allocation, the UAV perspective overlap rate and motion consistency constraints are used to establish a target consistency judgment model based on multi-perspective spatial overlap to determine whether the observed targets are the same target; Step 3: Design the state space, action space, and reward function based on the target consistency judgment model, and implement target capture collaborative decision-making through improved reinforcement learning; Step 4: Based on the target capture collaborative decision-making, a consensus drone cluster is formed based on group consensus, and the position and speed of the target are estimated by combining the observation error of the drone cluster and the target steering; Step 5: Based on the consensus UAV cluster’s estimation of the target’s position and velocity, a collaborative network graph relationship is constructed for the non-consensus UAVs through dynamic graph theory to perform spatiotemporal scheduling, optimize the allocation of capture points, and resolve conflicts.

2. The unmanned swarm distributed collaborative decision-making method for complex manhunt scenarios according to claim 1 is characterized by: The target consistency judgment model is: Among them, τ vol , τ area is the threshold of volume overlap ratio and projection overlap ratio; this constraint indicates that the view overlap needs to reach a certain level to be considered as a unified drone. is the distance difference between UAV j1 and UAV j2; the integral term reflects the target in the time window The maximum possible displacement within is the maximum flight speed of the UAV; the upper and lower limits of the time integration range 3. The unmanned swarm distributed collaborative decision-making method for complex search and arrest scenarios according to claim 2 is characterized by: The comprehensive hit rate model of drone cluster is: in, is the comprehensive hit rate.

4. The unmanned swarm distributed collaborative decision-making method for complex manhunt scenarios according to claim 3 is characterized by: Step 4 includes the following steps: Step 41: When the drone cluster recognizes the target at the same time, in order to reach a consensus, the consensus conditions of the drones that recognize the target are determined. First, the observation state of each drone i is expressed as follows: Among them, the state is the position vector, velocity vector, and angle of drone j in turn; If there are N identify UAVs recognize the target, and each UAV has its own recognition confidence p j , then for the cluster, the target is first judged by the consistency observation model, and then the spatiotemporal joint perspective overlap confidence ζ is obtained TS : Among them, N identify is the number of drones that identified the target in the cluster; is the recognition probability of each drone in the drone cluster; μ TS is the time decay coefficient, which controls the weight of historical data; The time difference between any two UAVs j1 and j2 in identifying the target; by introducing the time decay term Reduce the confidence level of asynchronous observations to avoid interference from stale data; The group reaches consensus if the spatiotemporal joint overlap confidence satisfies the following conditions: g TS >g th Among them, th is the consensus threshold; After forming a consensus cluster, the observation angle of each drone is set to a ray pointing from the drone position to the target, and the equation is: Step 42, optimizing the target position estimate using weighted least squares method; First, define the observation residual of UAV j That is, the vertical distance from the actual position of the target to the jth observation ray, which is used to measure the error of the observation data: It can be seen that the smaller the residual, the closer the actual position of the target is to the observation ray of UAV j; so we can get: in, is the observation weight, which is related to the observation accuracy; To solve this equation, we need to take the partial derivative and set the derivative to zero to obtain a system of linear equations: Arranged into matrix form, we can get: Among them are: Therefore, the predicted position of the target can be obtained, and the solution is: Step 43, calculate the target's velocity: Assuming that both the drone and the target are moving at a constant speed, the target's apparent velocity vector relative to drone i is for: in, is the target k velocity vector; Step 44: The UAV continuously observes the target. When the target rotates, the change in azimuth angle is determined by the velocity component v perpendicular to the line of sight. ⊥ Caused by the unit vector n perpendicular to the line of sight j for: Therefore, when UAV j observes the target, the target's azimuth angle change rate is for: in, is the distance between UAV j and target k; the azimuth angle change rate quantifies the impact of target motion on the observation angle and provides a basis for velocity estimation; Step 45, As unknowns, construct a system of linear equations: Similarly, writing it in matrix form: Among them are: Since the above matrix C observe It is not a square matrix, so the least squares solution is: Among them, W observe is the diagonal weight matrix; In step 46, after the drone cluster reaches consensus, it completes the collaborative prediction of the target through group consensus. The position vector and velocity vector of the target are as follows: in, represents the position vector of the target, Represents the velocity vector.

5. The unmanned swarm distributed collaborative decision-making method for complex search and arrest scenarios according to claim 4 is characterized by: Step 3 includes starting to establish a collaborative hunting method based on multi-agent deep reinforcement learning based on the target consistency judgment model. Assume that UAV j1 detects target k during the search process and maintains the same speed as the target after obtaining relevant detection information. During the flight, continuous group broadcasting is carried out. The broadcast information package is as follows in, is the ID number of drone j1; The time when UAV j1 detects the target; The position of UAV j1 when it detects the target; The recognition probability of the target detected by UAV j1; The moving direction of UAV j1 when it detects the target; is the real-time speed of UAV j1; is the number of assists provided by UAV j1; For the idle drone j2, it receives the information packet broadcast by drone j1 After that, the information in the information packet is processed to obtain in, is the spatiotemporal alignment distance between UAV j1 and UAV j2; is the moving speed of target k; is the moving direction of target k; is the position of target k; is the recognition probability of the target by UAV j1 after processing; in, is the information packet transmission delay between UAV j1 and UAV j2; d safe is the safe tracking distance of the drone; α fusion , β fusion are adjustment factors for speed and distance, which determine their impact on the recognition probability; and are calculated as follows: in, The time when drone j2 receives the information packet.

6. The unmanned swarm distributed collaborative decision-making method for complex search and arrest scenarios according to claim 5 is characterized by: Step 3 involves getting the packet Finally, we improve reinforcement learning: First, design its state space S, where the state vector s t for: The action vector a in the action space A t for: Among them, pos desired is the expected capture position; v desired is the output speed; θ desired is the output flight angle; t desired is the expected arrival time; is the assistance intensity to UAV j1, representing the willingness to assist; The reward function reward is designed as: in, The value is 1 when the capture is successful, otherwise it is 0; pos pred To predict the capture location, t pred is the predicted arrival time; α r , β r , γ r , δ r are weight coefficients respectively. For position matching and time matching, the higher the assistance intensity, the stricter the matching accuracy requirement. For assistance intensity rewards, it is encouraged to increase the assistance intensity moderately to avoid excessive dominance by a single drone. For each predictor variable in the reward function, the following requirements must be met: Among them, N request The number of drones making the request; Create a drone request list based on the assistance intensity assist The drones in the list are arranged in descending order of assistance strength: List assist ={j|φ j >φ j+1 },j=1,2…,N request Therefore, UAV j2 has the following choices: Among them, N request is the number of drones that send information packets; N assist The number of assisting drones needed for the drone to capture the target; assist means assisting the current drone j; search means searching operation; The drone that decides to assist in the roundup responds to the requesting drone. After receiving the response, the requesting drone increases the number of assists: in, The current number of assisted drones that have been assisted; after the required number of roundups is met, the drone is requested to stop broadcasting information packets; When multiple drones identify a target, they first use the consistency determination model to determine whether they have identified the same target, and then group the drones that have identified the same target: Among them, group m ,m∈[1,…,N goal ] is the UAV group m that recognizes the same target; N goal is the target number; The number of assists for each drone in the group remains the same at all times, namely: At the same time, the initial assistance quantity of each drone is: Among them, num(group m ) is group m the number of drones in the Each drone in the group broadcasts an information packet. After the other idle drones receive it, since each idle drone receives information packets from multiple drones, they calculate the strength of each drone sending an information packet and compare them according to the above decision. They select the one with the greatest strength to assist in the roundup and send a response at the same time.

7. The unmanned swarm distributed collaborative decision-making method for complex manhunt scenarios according to claim 6 is characterized by: Step 5 includes: the consensus drone cluster determines the target's position and speed through group consensus, and sends the target's position and speed to the remaining non-consensus drones. Then, the spatiotemporal relationship of the drone cluster is modeled based on dynamic graph theory. Through time difference compensation and spatial path optimization, the non-consensus drones can achieve collaborative capture of the target. Step 51, construct graph G capture : G capture ={V capture ,he capture } Among them, V capture is the node set, representing all drone nodes; ε capture is an edge set, representing the connection between nodes; the attributes of drone nodes are as follows: in, The consensus of each drone node in the drone cluster, including its own position own velocity vector and the time the message was sent Each drone node in the non-consensus drone cluster, including its own location own velocity vector and the time the information was received If UAV j1 and UAV j2 can communicate, then there is a communication edge, and the edge weight is the communication quality Among them, α comm is the communication quality quantization coefficient, is the distance between UAV j1 and UAV j2; At the same time, since there is a time difference when the non-consensus drone k receives the consensus information, the time difference needs to be compensated to predict the current position of the target: in, Predicting the location of the target; is the component of the target velocity along the x-axis and y-axis; Step 52: Generate candidate capture points. First, determine the safe capture area with the predicted target position as the center and a radius of r. safe Generate a uniformly distributed set of candidate points P capture : Step 53: Next, according to the target's moving direction, the priority of the candidate capture points is adjusted, and the drone is preferentially assigned to the candidate point ahead of the target's moving direction to improve the interception success rate: in, Candidate roundup sites; is the weight of the candidate capture point; is the candidate capture point (x capture ,y capture ) relative to the predicted target position The direction angle and the target moving direction angle are: Step 54, constructing a virtual collaborative network, firstly Join the space-time graph G capture , add a virtual edge between each non-consensus drone j2 and the candidate capture point, and the edge weight is the spatiotemporal collaboration cost in, are the current position and candidate capture point of UAV k respectively; α capture , β capture , γ capture , δ capture is the weight coefficient, satisfying α capture +β capture +γ capture +δ capture =1; the first path time term in the space-time coordination cost Arrival of drone j2 at the capture point The time cost of the second item is to encourage the selection of UAVs with short distance and fast speed; the second item is the time lag item. Penalize drones that receive information late and prioritize drones with updated data; the third speed matching item It indicates that the drone with a speed close to the target is easier to intercept; the fourth item is the candidate weight item This indicates that the collaborative cost of high-weight candidate points is reduced, and thus they are preferred; Each drone selects the local optimal capture point in, is the heuristic function, and we have: Among them, λ crowded is the congestion penalty coefficient; Indicates the congestion of the capture point and counts the currently selected capture point The number of drones should be controlled to avoid drone gathering as much as possible; Step 55: After pre-allocation, if multiple drones still choose the same capture point The auction mechanism is triggered; the conflicting drones broadcast bids The drone with the highest bid wins the roundup spot UAVs that fail the auction immediately fall back to the suboptimal candidate point according to the heuristic function sorting and recalculate the spatiotemporal coordination cost. If the suboptimal point is still occupied, a local broadcast is triggered to notify other UAVs to adjust their paths to avoid cascading conflicts. At the same time, considering that distance affects communication quality and makes it impossible to send information accurately, constraints are required when allocating points: in, To ensure basic communication quality; N capture is the number of drones participating in the roundup; it is necessary to ensure that when drone j is allocated, there is at least one consensus drone around it that maintains communication with it.

8. The unmanned swarm distributed collaborative decision-making method for complex search and arrest scenarios according to claim 7 is characterized by: The step 2 comprises the following steps: Step 21: In a complex search and rescue scenario, when multiple drones identify a target at the same time, the sensor view angle is modeled, and the view coverage area of each drone j is Defined as a 3D cone: in, Detect distance for drones; To calculate the view overlap rate, we define the view overlap area between UAV j1 and UAV j2 is the intersection of the two 3D cones: The available overlap metric is volume overlap ratio Projection overlap ratio Where Vol(·) is the volume, which represents the consistency of the overlap of the three-dimensional coverage area; Area(Proj(·)) is the area of the two-dimensional plane projection of the region, which represents the consistency of the two-dimensional plane coverage overlap; At the same time, if there is an observation time difference Δt between the two drones op : in, are the times when UAV j1 and UAV j2 observe the target respectively; The volume overlap ratio after attenuation is Projection overlap ratio for: Among them, μ op is the attenuation coefficient; Step 22: Based on step 21, a target consistency determination model can be obtained.

9. The unmanned swarm distributed collaborative decision-making method for complex search and arrest scenarios according to claim 8 is characterized by: The step 1 comprises the following steps: Step 11: Model the forest search area as a three-dimensional space-time grid G: G={(x,y,z,t)|x∈[1,M G ],y∈[1,N G ],z∈[1,H G ],t∈[1,T G ]} Among them, (x, y, z, t) are the x-coordinate, y-coordinate, z-coordinate and time step of the search area respectively; M G 、N G 、H G They are the maximum values in the x-coordinate direction, y-coordinate direction, and z-coordinate direction of the region, and the total number of two-dimensional grids is K G =M G ×N G ;T G is the maximum value of regional time; For a drone swarm performing a search and capture mission, it is represented as follows: Among them, N arrest is the number of drones in the drone swarm; And for each drone Contains the following properties: in, is the flight speed of the drone; The flight direction of the drone; is the current position of the drone; is the flight time of the drone; is the detection radius of the UAV; For the target in G, its specific position cannot be known. Set the target k to the speed and direction Move, then the state transition of the target can be modeled through the Markov process, and the prior probability distribution of drone j finding target k at node (x, y) can be obtained: in, Denotes the UAV j at time t searc The probability that node (x, y) finds target k; ) represents the UAV j at time t search -1 The probability that node (x′, y′) finds target k; The variance of the target position change reflects the target's motion uncertainty; At the same time, since the energy consumption limit of the UAV may affect its flight distance, a UAV energy consumption model is established: in, is the UAV flight energy consumption, E sense Sensing energy consumption for drones; and: in, is the current speed of the UAV; sensor is the sensing energy consumption, which is a fixed value; is the energy consumption coefficient, which is used to indicate the impact of this factor on energy consumption; Step 12: Define the total cost of UAV cluster collaborative search planning in, is the flight path length of the j-th UAV; is the flight time of the jth UAV; is the flight energy consumption of the j-th UAV search process; is the weight coefficient; Taking into account the target's movement and time decay, the comprehensive hit rate Defined as the probability that a cluster sees a target at least once: Step 13, to jointly minimize the total cost and maximize the comprehensive hit rate, it is necessary to minimize the objective function: in, Used to balance target dimensions; Constraints: in, The UAV at time t search coordinates of The UAV at time t search +1 coordinate; Δt search is the time step; is the grid area in G, is the coverage area of UAV j; is the indicator function. When the condition is met, the value of the indicator function is 1, otherwise it is 0; are the positions of UAVs j1 and j2 respectively; R comm is the shortest communication distance between drones; Step 14: Use simulated annealing algorithm to search for an approximate optimal solution to obtain an approximate optimal path allocation.

10. A system for implementing the unmanned swarm distributed collaborative decision-making method for complex search and arrest scenarios as described in claim 1, characterized by: It includes an optimal path unit, a target consistency judgment model unit, an improved reinforcement learning unit, a target position and speed estimation unit, and a collaborative network unit, among which: The optimal path unit is used to model the search area as a three-dimensional space-time grid, define the drone cluster attributes and target state transition probabilities, and then establish a drone energy consumption model and a drone cluster comprehensive hit rate model. The total cost function is constructed by combining the flight path, time and energy consumption constraints, and the objective function is minimized through a simulated annealing algorithm to solve the approximate optimal path allocation; The target consistency judgment model unit is used to establish a target consistency judgment model based on multi-view spatial overlap according to the obtained approximate optimal path allocation through the drone perspective overlap rate and motion consistency constraint, and judge whether the observed targets are the same target; The improved reinforcement learning unit is used to design the state space, action space and reward function based on the target consistency judgment model, and realize the target capture collaborative decision-making through improved reinforcement learning; The target position and speed estimation unit is used to form a consensus drone cluster based on group consensus according to the target capture collaborative decision, and estimate the position and speed of the target in combination with the unmanned cluster observation error and target steering; The collaborative network unit is used to estimate the position and speed of the target based on the consensus drone cluster, and to construct a collaborative network graph relationship for non-consensus drones through dynamic graph theory, so as to perform spatiotemporal scheduling, optimize the allocation of capture points and resolve conflicts.

Citation Information

Cited By

  • Heterogeneous aircraft cluster collaborative hunting search method oriented to denial environment

    CN120686875A

  • Heterogeneous aircraft cluster cooperative hunting search method for denial environment

    CN120686875B

  • Unmanned aerial vehicle cluster collaborative target hunting method based on brain-like calculation

    CN120848561A

  • Distributed optical storage network energy regulation and control system based on unmanned aerial vehicle cluster

    CN121238651A