Unmanned Aerial Vehicle Cluster Cooperative Rescue Task Allocation Method Based on Population Interaction Learning Particle Swarm Optimization Algorithm

Through the population interactive learning particle swarm algorithm, the drone cluster collaborative rescue task allocation model is constructed, which solves the problem of insufficient drone task allocation efficiency in post-disaster rescue, and realizes efficient allocation of rescue materials and time optimization, improving post-disaster rescue efficiency.

CN119443626BActive Publication Date: 2025-07-08XUZHOU NORMAL UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411501275.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-25
Publication Date
2025-07-08
Estimated Expiration
2044-10-25

AI Technical Summary

Technical Problem

In post-disaster rescue, the existing drone mission allocation methods are insufficient in the face of complex scenarios such as uncertainty and multi-target and multi-preference, and it is difficult to efficiently allocate rescue materials in the shortest time and ensure the safety of trapped people.

Method used

The particle swarm algorithm based on population interactive learning is adopted to construct the minimum rescue comprehensive punishment objective function and the shortest rescue time objective function. The coordinated rescue task allocation of the drone cluster is optimized through K-means clustering and population interactive learning particle swarm algorithm, and combined with integer encoding and decoding strategies to achieve efficient allocation and sorting of drone tasks.

Benefits of technology

It improves the efficiency of post-disaster rescue work, significantly accelerates the convergence speed of the population, optimizes the distribution of Pareto frontier, has good robustness and adaptability, and can stably solve the problem of collaborative multi-task allocation of drone clusters in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119443626B_ABST
    Figure CN119443626B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for collaborative rescue task allocation of UAV swarms based on a population interactive learning particle swarm algorithm, comprising the following steps: determining a minimum rescue comprehensive penalty objective function based on the rescue value of each rescue task and the time for each UAV to reach the rescue task it is responsible for from the base; determining a shortest rescue time objective function based on the time for each UAV to reach the last rescue task it is responsible for; constructing a collaborative rescue task allocation model for UAV swarms according to the minimum rescue comprehensive penalty objective function and the shortest rescue time objective function; solving the collaborative rescue task allocation model for UAV swarms based on the population interactive learning particle swarm algorithm, and allocating the collaborative rescue tasks of UAV swarms based on the solution results. This method can effectively solve the problem of collaborative multi-task allocation of UAV swarms to improve the efficiency of post-disaster rescue work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of UAV mission allocation, and more specifically, to a UAV swarm collaborative rescue mission allocation method based on a population interaction learning-based particle swarm algorithm. Background Art

[0002] In the traditional post-disaster rescue process, challenges such as insufficient information and safety risks are often faced. The introduction of unmanned aerial vehicles (UAVs) provides new possibilities for solving rescue work. The key to the post-disaster rescue problem is to deliver rescue supplies to as many trapped people as possible in the shortest time. Therefore, according to the pre-known disaster and UAV information, accurately planning the tasks and execution order of each UAV, that is, mission allocation, becomes the core link of rescue supply delivery. Therefore, reasonable mission allocation is the core to ensure the efficient execution of various tasks by UAVs. For the multi-UAV mission allocation problem, the above process includes two stages: allocation and sorting. Allocation refers to allocating multiple tasks to different UAVs; sorting is to sort the execution order of the allocated tasks.

[0003] It is easy to know that after a disaster occurs, the mission information is uncertain. Due to the interference caused by the collapse of buildings and the like, rescue personnel often cannot accurately obtain the number of trapped people, and it is difficult to effectively carry out rescue operations. Therefore, the research on UAV mission allocation strategies in an uncertain environment is of great significance.

[0004] In recent years, scholars have conducted a large number of studies on the mission allocation problems for different scenarios and UAV types, and proposed various UAV mission allocation models and corresponding algorithms applicable to different problems and scenarios. However, existing methods are mostly designed based on precise targets. For fuzzy and uncertain mission situations, due to the addition of uncertain factors, in scenarios such as multi-objective, multi-preference, and strong constraints, the traditional mission allocation methods are slightly insufficient in effectiveness.

[0005] The particle swarm optimization (PSO) algorithm has performed prominently in solving the UAV collaborative mission allocation problem in recent years due to its few parameters, insensitivity to parameters, and easy expansion, and there have been many successful application cases. Therefore, in an environment where mission information is uncertain, especially when facing different weights and time-sensitive rescue missions, how to use the particle swarm algorithm to effectively solve the multi-UAV collaborative multi-mission allocation problem to improve the efficiency of post-disaster rescue work is an urgent problem for those skilled in the art to solve. Summary of the Invention

[0006] In view of the above problems, the present invention provides a method for allocating cooperative rescue tasks for an unmanned aerial vehicle (UAV) cluster based on a population interaction learning-based particle swarm optimization algorithm, so as to solve at least some of the technical problems mentioned in the above background art.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] The present invention provides a method for allocating cooperative rescue tasks for an unmanned aerial vehicle (UAV) cluster based on a population interaction learning-based particle swarm optimization algorithm, including the following steps:

[0009] Based on the rescue value of each rescue task and the time for each UAV to reach the rescue task it is responsible for from the base, determine the minimum rescue comprehensive penalty objective function;

[0010] Based on the time for each UAV to reach the last rescue task it is responsible for, determine the shortest rescue time objective function;

[0011] According to the minimum rescue comprehensive penalty objective function and the shortest rescue time objective function, construct a cooperative rescue task allocation model for the UAV cluster;

[0012] Based on the population interaction learning-based particle swarm optimization algorithm, solve the cooperative rescue task allocation model for the UAV cluster, and allocate the cooperative rescue tasks for the UAV cluster based on the solution result.

[0013] Furthermore, the minimum rescue comprehensive penalty objective function is expressed as:

[0014]

[0015] Wherein, f1 represents the minimum rescue comprehensive penalty objective function; I represents a total of I UAVs; i represents the i-th UAV; J represents a total of J rescue tasks; j represents the j-th rescue task; represents the time for UAV i to reach rescue task j from the base; represents the rescue value of rescue task j at time t.

[0016] Furthermore, the corresponding obtaining steps for the time for each UAV to reach the rescue task it is responsible for from the base include:

[0017] Based on the maximum no-load flight speed of the UAV, the maximum load capacity of the UAV, and the amount of rescue supplies of each UAV before reaching each rescue task, obtain the flight speed of each UAV between every two rescue tasks; expressed as:

[0018]

[0019] Wherein, vel i,jDenote the flight speed of UAV i between rescue mission j - 1 and rescue mission j; rescue mission j - 1 is the previous rescue mission of rescue mission j; ε represents the UAV performance impact coefficient; vel max Denote the maximum no - load flight speed of the UAV; S max Denote the maximum load capacity of the UAV; S i,j Denote the amount of rescue supplies of UAV i before reaching rescue mission j;

[0020] Based on the flight speed and distance of each UAV between every two rescue missions, obtain the time for each UAV to reach the rescue missions it is responsible for from the base, expressed as:

[0021]

[0022] Among them, Denote the time for UAV i to reach rescue mission j from the base; J i Denote the number of rescue missions executed by UAV i from leaving the base to returning to the base, and Denote the distance between rescue mission j - 1 and rescue mission j; ΔT represents the sum of the unloading time and replenishment time of UAV i.

[0023] Furthermore, the shortest rescue time objective function is expressed as:

[0024]

[0025] Among them, f2 represents the shortest rescue time objective function; Denote the time for UAV i to reach the last rescue mission it is responsible for; end represents the last rescue mission; I represents there are I UAVs in total; ΔT represents the sum of the unloading time and replenishment time of UAV i.

[0026] Furthermore, the encoding and decoding strategies of the population interactive learning - based particle swarm algorithm specifically include:

[0027] In the encoding stage: Integer encoding is adopted, and the particle structure is a two - stage representation method; The first stage is a nested array, which is a matrix of size 1*I, indicating that there are I inner arrays. The i - th column array X 1i has an inner size of 1*n i , indicating the rescue missions assigned to UAV i and their execution order, n i is the number of rescue missions assigned to UAV i; Suppose there are J rescue missions in total, so the value range of X 1i is [1, J]; The second stage is an array X i of size 1*ns 2i , indicating the replenishment moment of UAV i, represented by the rescue mission label, that is, the UAV replenishes after completing this mission, nsi is the number of resupplies required during the mission of UAV i, and its value range is {1, 2, …, n i - 1};

[0028] In the decoding stage: a two - stage embedding method is adopted, that is, in the first stage, the starting point, resupply points and ending point are embedded according to the content of the array in the second stage; if decoding the i - th particle X i perform decoding, first decode the j - th column of X i in the first stage, separate the mission assignment of the i - th UAV, that is, a series of mission points, denoted as X i ; according to the resupply point position information given at the corresponding position in the second stage, embed the resupply points into the mission; the embedding operation follows the principle of proximity, that is, select the resupply point closest to the first rescue mission to be executed after the UAV is resupplied, and finally embed the base at both ends of the UAV mission assignment sequence, thus obtaining the true flight path of the decoded UAV, and this flight path includes the starting point, rescue points, resupply points and ending point.

[0029] Furthermore, the population - interaction - learning - based particle swarm optimization algorithm is used to solve the UAV swarm cooperative rescue mission assignment model, specifically including:

[0030] Initialize the population - interaction - learning - based particle swarm optimization algorithm based on K - means clustering to obtain the initial population;

[0031] Divide the initial population into the main population and the auxiliary population;

[0032] Determine the particles to be updated in the main population, and select particle learning objects based on the auxiliary population;

[0033] Update the particles to be updated based on the particle learning objects.

[0034] Furthermore, the initializing the population - interaction - learning - based particle swarm optimization algorithm based on K - means clustering to obtain the initial population specifically includes:

[0035] (1) According to the position information of J rescue missions, use the K - means clustering algorithm to divide these J rescue missions into I sub - clusters; randomly assign the I sub - clusters to I UAVs, thus completing the initial mission assignment;

[0036] (2) Based on the comprehensive benefits corresponding to the rescue missions responsible for each UAV, sort the rescue missions of each UAV; specifically: assume that rescue mission j is one of the rescue missions responsible for UAV i, and the comprehensive benefit of selecting rescue mission j as the next rescue mission to be executed by UAV i is:

[0037]

[0038] Among them, represents the comprehensive benefit of the UAV i selecting the rescue task j as the next rescue task at time t; v j (t) represents the rescue value of the rescue task j; represents the time for the UAV i to reach the rescue task j from the base; d i,j represents the distance between the rescue task j and the UAV i; α represents the weight;

[0039] (3) Under the current rescue task execution order of the UAV i, replenish the UAV; specifically: assume that after the UAV i completes the jth rescue task assigned, the total amount of rescue supplies unloaded by the UAV i is At this time, the replenishment probability of the UAV i is expressed as:

[0040]

[0041] Among them, represents the replenishment probability of the UAV i after completing the jth rescue task; S max represents the maximum load capacity of the UAV; In particular, when and it means that the UAV i cannot complete the jth rescue task assigned. In this case, after completing the (j - 1)th rescue task, the UAV must be replenished.

[0042] Furthermore, the division of the initial population into the main population and the auxiliary population specifically includes:

[0043] Perform non - dominated sorting on the initial population Pop in the form of the number of intervals to obtain a new population Pop' with a new particle sorting;

[0044] Determine the smallest integer θ divided according to the principle of the least common multiple according to the ratio θ int ;

[0045] According to the non - dominated sorting result, select the first θ int particles from the new population Pop' to form a smallest cycle unit, and distribute the particles to the main population Pop1 and the auxiliary population Pop2 within the smallest cycle unit according to the ratio θ; continue to select the next θ int particles for division according to the non - dominated sorting result, and repeat this process until all the particles in the new population Pop' are distributed.

[0046] Furthermore, the determination of the particle learning object based on the main population and the auxiliary population specifically includes:

[0047] Take the particles in the auxiliary population that are not dominated by the particle to be updated as the range of the particle learning object;

[0048] Select the particle with the shortest vertical distance between the line connecting the particle to be updated and the origin within the range as the particle learning object.

[0049] According to the above technical solutions, compared with the prior art, the present invention discloses a method for allocating UAV swarm cooperative rescue tasks based on a population interaction learning particle swarm algorithm, which has the following beneficial effects:

[0050] The present invention takes the minimum rescue penalty and the minimum task completion time as evaluation indicators, and establishes a UAV swarm cooperative rescue task allocation model under uncertain environments. This model can effectively solve the problem of UAV swarm cooperative multi-task allocation to improve the efficiency of post-disaster rescue work.

[0051] The population interaction learning particle swarm algorithm provided by the present invention significantly accelerates the convergence speed of the population and optimizes the distribution of the Pareto front through parallel update of the main and auxiliary populations, information interaction, and collaborative evolution. This algorithm can not only stably solve the problem of UAV swarm cooperative task allocation, but also has good robustness and stability, and shows better adaptability in complex scenarios.

[0052] The technical solutions of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0054] Figure 1 It is a schematic flowchart of a method for allocating UAV swarm cooperative rescue tasks based on a population interaction learning particle swarm algorithm provided by an embodiment of the present invention.

[0055] Figure 2 It is a schematic diagram of rescue task allocation provided by an embodiment of the present invention.

[0056] Figure 3 It is a schematic diagram of the framework of a population interaction learning particle swarm algorithm provided by an embodiment of the present invention.

[0057] Figure 4 It is a schematic diagram of initial task allocation provided by an embodiment of the present invention.

[0058] Figure 5 It is a schematic diagram of population division provided by an embodiment of the present invention.

[0059] Figure 6Schematic diagram of learning object selection provided by an embodiment of the present invention.

[0060] Figure 7 Schematic diagram of learning object selection proof provided by an embodiment of the present invention.

[0061] Figure 8 Schematic diagram of two-stage update provided by an embodiment of the present invention.

[0062] Figure 9 Schematic diagram of internal population update of the learning library provided by an embodiment of the present invention.

[0063] Figure 10 Schematic diagram of evaluation index provided by an embodiment of the present invention.

[0064] Figure 11(a) is a box diagram obtained by the PILPSO provided by an embodiment of the present invention and four comparison algorithms in Scenario 1.

[0065] Figure 11(b) is a box diagram obtained by the PILPSO provided by an embodiment of the present invention and four comparison algorithms in Scenario 2.

[0066] Figure 11(c) is a box diagram obtained by the PILPSO provided by an embodiment of the present invention and four comparison algorithms in Scenario 3.

[0067] Figure 12(a) is a schematic diagram of the Pareto front distribution of five algorithms provided by an embodiment of the present invention in Scenario 1.

[0068] Figure 12(b) is a schematic diagram of the Pareto front distribution of five algorithms provided by an embodiment of the present invention in Scenario 2.

[0069] Figure 12(c) is a schematic diagram of the Pareto front distribution of five algorithms provided by an embodiment of the present invention in Scenario 3. Detailed implementation manners

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0071] Refer to Figure 1 As shown, an embodiment of the present invention discloses a method for allocating cooperative rescue tasks for an unmanned aerial vehicle (UAV) cluster based on a population interaction learning-based particle swarm optimization algorithm, including the following steps:

[0072] Determine a minimum rescue comprehensive penalty objective function based on the rescue value of each rescue task and the time for each UAV to reach the responsible rescue task from the base.

[0073] Determine the shortest rescue time objective function based on the time when each drone reaches the last rescue task it is responsible for;

[0074] Construct a cooperative rescue mission allocation model for the drone swarm according to the minimum rescue comprehensive penalty objective function and the shortest rescue time objective function;

[0075] Solve the cooperative rescue mission allocation model for the drone swarm based on the population interactive learning particle swarm algorithm, and allocate the cooperative rescue missions for the drone swarm based on the solution results.

[0076] Next, the above content will be described in detail.

[0077] 1. Problem Description and Model Establishment

[0078] 1.1 Problem Description

[0079] To simplify the problem-solving process, without loss of generality, the following assumptions are made:

[0080] 1) When the drone executes tasks under a fixed load, its flight speed is constant;

[0081] 2) Before the rescue, use the drone to detect the vital signs of the environment and determine the location of each rescue point;

[0082] 3) During the entire mission execution process, the positions of the trapped people at all rescue points remain unchanged;

[0083] 4) The material supply station (i.e., the supply point) has sufficient resources to ensure the completion of the mission.

[0084] Based on the above assumptions, in a certain post-disaster environment, assume there are I drones parked at the same base, and each drone i corresponds to a rescue mission list O i ; m supply points, scattered at several accessible fixed points within the area, and their positions are known; there are a total of J rescue points forming a rescue mission set [1, J], and the positions of the rescue points are known, expressed as P = {(x1, y1), (x2, y2), …, (x J , y J )}. The drones have certain limitations in terms of flight range, cargo capacity, and speed. It is required that the drones cooperate with each other. Under the condition of uncertain mission information, calculate the rescue value information at different times, and based on this, plan the mission execution sequence for each drone, in order to complete the mission with the least time under the premise of the minimum rescue penalty.

[0085] A schematic diagram of multi-drone cooperative rescue can be seen in Figure 2As shown in the figure, there are 3 drones in the environment; 12 rescue tasks are required, each rescue task corresponding to a rescue point; and two supply points. Figure 2 In the figure, the arrows of different colors indicate the rescue paths of different aircraft. Taking the rescue path of UAV 1 as an example, the order of its task rescue is successively rescue point 1 → supply station → rescue point 8 → rescue point 2 → rescue point 4.

[0086] 1.2 Calculation of rescue value

[0087] In order to complete the tasks with the minimum rescue penalty in the shortest time, it is necessary to judge the rescue value according to the specific situation of each task. The rescue value is usually related to the number of rescued people, the rescue difficulty and the rescue time, but the number of rescued people can more directly judge the value of a rescue point. After a disaster occurs, due to limited high-altitude reconnaissance information, only the approximate number range of trapped people can be judged through the building scale, damage degree and other pre-known information. Therefore, in the embodiments of the present invention, based on the building damage degree map obtained by satellites or high-altitude drones, and based on the model and results constructed in the "Earthquake Casualty Assessment Method and Application Based on Building Structure Damage" published in the Journal of Tsinghua University (Science and Technology Edition) in 2015 (hereinafter referred to as Document 1), the damage situation of the building where the rescue point is located is judged and inferred, and the present invention calculates the number of trapped people in the rescue point based on this result.

[0088] In the embodiments of the present invention, the rescue priorities are divided into three levels: serious (H), medium (M) and light (L) injuries, and the number of wounded in each rescue point is determined according to prior information. Considering the complexity and uncertainty of the post-disaster environment, the number of people calculated in Document 1 is not accurate enough. Therefore, in the embodiments of the present invention, interval numbers are used to represent the number of wounded in each level. Specifically, assuming that the calculation result based on the model in Document 1 is Pr, then the interval number of the present invention is expressed as [Pr(1 - err), Pr(1 + err)], where err is the ratio of the average error between the calculation result of the model in Document 1 and the actual result. From the results of Document 1, it can be seen that err for the number of lightly wounded is 27%, err for the number of moderately wounded is 20%, and err for the number of seriously wounded is 12%. Then, the present invention determines the number interval of each wounded level based on this.

[0089] When setting the function, taking the sum of the number intervals of all single-injury-degree wounded in the entire area to be rescued as the benchmark, and using the ratio of the number intervals of wounded of different degrees in different rescue points to the total sum as the calculation basis, the absolute number is converted into a relative number to make it more in line with the actual situation, as shown in Equation (1):

[0090]

[0091] Where and respectively represent the number of intervals of seriously injured, moderately injured, and slightly injured people in rescue point j; the total number of seriously injured people is moderately injured is slightly injured is

[0092] During the rescue process, since the condition of the wounded will deteriorate over time, especially for moderately and seriously injured patients, the rescue value of the rescue point will also increase with time. As time goes by, in the rescue value and parts will increase, which can be calculated by Equations (2-1) and (2-2):

[0093]

[0094]

[0095] Among them, γ1 and γ2 represent the growth coefficients, and their values are set according to the actual situation. t is the time, and the finally obtained rescue value is a 3×2 matrix; v h (t), v m (t) and v l (t) respectively represent the rescue values of seriously injured, moderately injured, and slightly injured; the rescue value weights for each injury level from serious to minor are ω1, ω2, and ω3, and the values of the weights are determined by the decision maker according to the actual situation; and Similarly.

[0096] 1.3. Uncertain optimization model

[0097] The embodiment of the present invention considers two objective functions: the minimum rescue comprehensive penalty objective function f1 and the shortest rescue time objective function f2. In f1, priority is given to tasks with higher rescue values, and the integral value of the growth curve of its rescue value over time is used as the rescue penalty. According to Section 1.2, the rescue value is the number of intervals, so f1 is constructed as an interval function. The other objective function f2 aims to complete all rescue tasks in the shortest time. Since these two objectives conflict, the embodiment of the present invention uses two objective functions to provide different options for the decision maker to choose from.

[0098] In model construction, the influence of the load capacity of rescue supplies on the flight speed of the UAV also needs to be considered. In addition to the UAV base, multiple resource supply points are set in the area for the UAV to replenish supplies. Therefore, the decision variables of the model constructed in the embodiment of the present invention are task allocation, task scheduling after allocation, and the choice of replenishment time.

[0099] 1.3.1. Objective function

[0100] (1) The minimum rescue comprehensive penalty objective function f1 is expressed as:

[0101]

[0102] Among them, f1 represents the minimum rescue comprehensive penalty objective function; I represents a total of I drones; i represents the i-th drone; J represents a total of J rescue tasks; j represents the j-th rescue task; represents the time for the drone i to reach the rescue task j from the base; represents the rescue value of the rescue task j at time t (the rescue values of the base and the supply station are both 0), and this value is calculated by Equation (2) and is a 3×2 matrix; represents the cumulative rescue penalty of the drone i reaching the rescue task j for trapped persons with different injury degrees, and it is also a 3×2 matrix; represents the time for the drone i to reach the rescue task j from the base; its calculation method is shown in Equation (4).

[0103]

[0104] Among them, represents the time for the drone i to reach the rescue task j from the base; J i represents the number of rescue tasks executed by the drone i from leaving the base to returning to the base, and represents the distance between the rescue task j - 1 and the rescue task j; ΔT represents the sum of the unloading time and the replenishment time of the drone i; vel i,j represents the flight speed of the drone i between the rescue task j - 1 and the rescue task j, and its value is related to its load capacity, and is specifically expressed as shown in Equation (5).

[0105]

[0106] Among them, vel max represents the maximum no-load flight speed of the drone; S max represents the maximum load capacity of the drone; S i,j represents the amount of rescue supplies of the drone i before reaching the rescue task R ,j before; ε represents the drone performance influence coefficient, and its value is determined according to the drone performance.

[0107] (2) The shortest rescue time objective function f2

[0108] The rescue time f2 of the drone to complete all rescue tasks can be calculated by Equation (6):

[0109]

[0110] Among them, For the moment when the drone i reaches the last rescue mission point in the mission execution list O i ; end represents the last rescue mission; ΔT represents the sum of the unloading time and replenishment time of the drone i.

[0111] 1.3.2. Constraints

[0112] During the mission execution, the drone needs to consider multiple constraints, including battery and payload constraints, etc. The specific requirements are as follows:

[0113] (1) Drone travel power consumption constraint, that is, the power consumption of each drone from the base to returning to the base should be less than or equal to the drone's battery capacity; expressed as:

[0114]

[0115] Among them, J i represents the number of rescue missions executed by the i-th drone from leaving the base to returning to the base. C b represents the battery capacity of the drone; ρ represents the distance that can be flown per unit of power;

[0116] (2) Drone load constraint, that is, the total amount of supplies unloaded by each drone after leaving the current base or supply point should be less than or equal to the maximum cargo capacity of the drone; expressed as:

[0117]

[0118] Among them, O i represents the rescue mission list corresponding to the drone i; J i ’ represents the consecutive rescue missions between any two bases / supply stations in O i ; for example, if O i = [0,1,5,1002,7,8,0], then J i ’ = [1,5] or J i ’ = [7,8], where 102 represents returning to the supply station and 0 represents the base station; represents the amount of rescue supplies corresponding to J i ’; S max represents the maximum cargo capacity of the drone;

[0119] (3) Task assignment constraint, that is, each rescue mission can only be assigned to one drone, and each rescue mission must be assigned to a drone for execution; expressed as:

[0120]

[0121] Among them, represents the sum of the number of all mission points in the j-th column of the matrix x; Denote the number of task \(j\) in the \(i\)-th column of \(x\). \(\delta()\) represents the indicator function, where \(\delta() = 1\) when the condition inside the parentheses is true; otherwise, it is 0.

[0122] (4) UAV resupply constraint, that is, when a UAV completes a rescue task and the remaining rescue supplies do not meet the requirements of the next rescue task, it needs to return to the nearest supply point to replenish supplies; expressed as:

[0123]

[0124] Among them, \(O\) i,j+1 represents the \((j + 1)\)-th rescue task in the rescue task list \(o\) corresponding to UAV \(i\); \(O\) i represents the \(j\)-th rescue task in the rescue task list \(O\) corresponding to UAV \(i\); i,j represents the remaining rescue supplies when UAV \(i\) completes \(O\) i ; \(r_{i,j}^o\) represents the amount of rescue supplies required when UAV \(i\) completes \(O\) represents the remaining rescue supplies when UAV \(i\) completes \(O\) i,j ; \(r_{i,j}^o\) represents the amount of rescue supplies required when UAV \(i\) completes \(O\) represents the remaining rescue supplies when UAV \(i\) completes \(O\) i,j+1 ; \(r_{i,j}^o\) represents the amount of rescue supplies required when UAV \(i\) completes \(O\). \(M\) represents an integer, and in this invention, it takes the value of 1000; \(m\) represents the number of supply points;

[0125] (5) UAV path constraint, that is, each UAV needs to start from the base at the beginning and return to the base after completing the assigned rescue tasks; expressed as:

[0126] \(O_{i,1}^o = 0\) (11) i,1 \(O_{i,J_i^o}^o = 0\) (12)

[0127] Among them, \(O_{i,1}^o\) i,end represents the first rescue task responsible for by UAV \(i\); \(O_{i,J_i^o}^o\)

[0128] represents the last rescue task responsible for by UAV \(i\); i,1 represents the first rescue task responsible for by UAV \(i\); \(O_{i,J_i^o}^o\) i,end represents the last rescue task responsible for by UAV \(i\);

[0129] (6) UAV task volume constraint, that is, each UAV executes at least one task. Expressed as:

[0130]

[0131] Among them, \(J_i^o\) i represents the number of rescue tasks executed by the \(i\)-th UAV from leaving the base to returning to the base.

[0132] 1.4. Model transformation

[0133] The value of the minimum rescue comprehensive penalty objective function f1 in the model established in the previous section is an interval number. For the convenience of calculation, this section will convert the interval objective function into an exact objective function. In the conversion process, in addition to considering the deterministic information of the interval number, its uncertainty information also needs to be considered. Therefore, this section will consider both the interval weighted mean and the interval uncertainty degree to convert the interval objective value into an exact value. Specifically, by setting the decision maker's tolerance coefficient ζ∈[0,1], the minimum rescue comprehensive penalty objective function f1 is converted into a deterministic objective, as shown in formula (14).

[0134]

[0135] Among them, Punish u is used to evaluate the uncertainty degree of the interval minimum rescue comprehensive penalty objective function f1, and its value is Punish + -Punish - ; Punish m is the interval weighted mean that integrates the decision maker's preference, Punish m = Punish - + Punish u ×ζ. If the decision maker is more optimistic, the coefficient ζ can take a smaller value; on the contrary, the value of ζ can be increased. Using Punish u and Punish m as the horizontal and vertical coordinates respectively to determine the position of the original f1 after conversion on the plane, and based on formula (14), convert the position information of this interval number on the plane into the vector length Punish exact , and replace the fuzzy interval information with the exact vector length information.

[0136] In addition, in the algorithm proposed in the present invention, to reduce the influence of information distortion caused by conversion, the objective function value represented by the interval number is used for the comparison and non-dominated sorting operations of the main population. Therefore, an interval comparison method is defined: the weighted two-end subtraction comparison method: let the interval number a = [a - , a + , the interval number b = [b - , b + ; when ζ×(b + - a + )+(1 - ζ)×(b - - a - )> 0, then b > a, otherwise a > b; if ζ×(b + - a + )+(1 - ζ)×(b - - a - ) = 0, then b = a.

[0137] 2. Algorithm Design

[0138] To efficiently solve the above model, the present invention proposes a population interaction learning-based particle swarm optimization algorithm. Figure 3 The basic process framework of the proposed algorithm is given. First, the population is initialized according to the characteristics of the problem, and the population is non-dominated sorted. Then, according to the given ratio θ and the population sorting result, the population is divided into a main population and an auxiliary population. Among them, the main population is responsible for the iteration and convergence of the algorithm, and is supplemented by local update to optimize the distribution of solutions; the auxiliary population serves as an evolvable external archive set to provide guidance for the main population. The two populations run in parallel and exchange some individuals regularly. The two populations adopt different strategies to execute the basic PSO process to update the individual and global extreme values; at the same time, the main population performs non-dominated sorting in the form of the number of intervals on the population at the end of each generation to reduce the impact of information distortion caused by numerical conversion on the reliability of the solutions.

[0139] Specifically, in the "problem-guided population initialization strategy" stage, according to the problem characteristics, strategies are gradually selected to initialize it to obtain a set of better initial population distributions; in the "update learning based on population interaction" stage, the iteration update of the main population is guided by constructing an auxiliary population, focusing on accelerating the algorithm convergence and optimizing the diversity of the Pareto front. The present invention constructs a learning library with the auxiliary population, updates it in parallel with the main population, and at the same time serves as the learning and interaction object of the main task.

[0140] The following Table 1 gives the pseudocode of the population interaction learning-based particle swarm optimization algorithm proposed by the present invention. The first line executes the population initialization strategy in Section 2.2, and initializes the allocation, sorting, and supply time selection in sequence; the second line executes the population splitting operation before the iterative update operation, and divides the population Pop into the main population Pop1 and the auxiliary population Pop2 according to the given ratio θ and the population sorting result; the third to eleventh lines are the iterative update stage, where the fifth to sixth lines are the internal self-update stage of Pop2; the seventh line is the operation of selection, update of Pop1 and local optimization according to the sparse area of the front; the eighth to ninth lines are the individual exchange between Pop1 and Pop2 according to the conditions; the tenth line is the subsequent processing operation of each generation.

[0141] The pseudocode of the proposed algorithm.

[0142] Table 1 Overall operation process of the swarm interaction learning-based particle swarm optimization algorithm

[0143] Input Population size Num, maximum number of iterations G Output Optimal solution (1) Initialize PSO parameters, determine task assignment, task sequencing, and replenishment time (2) <![CDATA[Sort the population Pop and split it into the main population Pop1 and the auxiliary population Pop2;]]> (3) <![CDATA[Convert the numerical forms of Pop1 and Pop2;]]> (4) for ger = 1:G % ger is the current iteration number (5) <![CDATA[Auxiliary population Pop2 update: Random strategy update;]]> (6) <![CDATA[Update of main population Pop1: Select learning objects from Pop2 and guide the particles to update their positions;]]> (7) Optimize the distribution of the Pareto front; (8) <![CDATA[If the interaction condition between Pop2 and Pop1 is satisfied, exchange some individuals between Pop2 and Pop1;]]> (9) <![CDATA[Sort Pop1 and Pop1;]]> (10) <![CDATA[Perform numerical form conversion and non-dominated sorting on Pop1;]]> (11) end (12) Output the optimal solution

[0144] 2.1. Encoding and decoding

[0145] The embodiments of the present invention aim to optimize three objective variables: task allocation, the task scheduling sequence after allocation, and the replenishment time. The PSO algorithm is used for solving, where a particle represents a solution to the problem. To simplify the decoding process, integer encoding is adopted, and the particle structure is a two-stage representation method. The first stage is a nested array, which is a 1*I matrix, indicating that there are I inner arrays. The size of the i-th column array X 1i is 1*n i , representing the rescue tasks assigned to the UAV i and their execution order. n i is the number of rescue tasks assigned to the UAV i. Suppose there are J rescue tasks in total. Therefore, the value range of X 1i is [1, J]; the second stage is an array X i of 1*ns 2i , representing the replenishment time of the UAV i, which is represented by the rescue task label, that is, the UAV replenishes after completing this task. ns i is the number of times the UAV i needs to replenish during the mission, and its value range is {1, 2, …, n i -1};

[0146] In the decoding stage: a two-stage embedding method is adopted, that is, in the first stage, the starting point, replenishment point, and ending point are embedded according to the content of the array in the second stage. Suppose decoding the i-th particle X i . First, decode the j-th column of the first stage of X i to separate the task allocation situation of the i-th UAV, that is, a series of task points, represented as X i ; according to the replenishment point position information given at the corresponding position in the second stage, embed the replenishment point after the task. The embedded element is a relatively large integer M (the element embedded in the present invention is the replenishment point serial number plus 1000) to show the difference from the rescue task. The embedding operation follows the principle of proximity, that is, select the replenishment point closest to the first rescue task to be executed after the UAV replenishes. Finally, embed the base (represented by 0) at both ends of the UAV task allocation sequence, and thus obtain the true flight path of the decoded UAV. This path includes the starting point, rescue points, replenishment points, and ending point.

[0147] Taking the particle shown in Table 2 as an example, it shows the scenario where 3 UAVs execute 12 rescue tasks. Suppose there are two replenishment points in this scenario, represented as 1001 and 1002. If the rescue tasks assigned to the UAV 1 are [2, 5, 9, 6], and it needs to replenish after completing the second rescue task. The replenishment point closer to task 5 is 1002, and select it as the replenishment point. Therefore, its path is O1 = [0, 2, 5, 1002, 9, 6, 0]; similarly, O2 = [0, 8, 4, 1001, 5, 7, 1002, 12, 0], O3 = [0, 11, 10, 0].

[0148] Table 2 Particle Two-Stage Coding Examples

[0149] UAV number <![CDATA[UAV1]]> <![CDATA[UAV2]]> <![CDATA[UAV3]]> Assigned tasks (first stage) [2,5,9,6] [8,4,1,5,7,12] [11,10] Replenishment time (second stage) [2] [2,5] [\]

[0150] In addition, regarding the specific amount of supplies replenished when the UAV is at the supply station, the maximum value required to complete the rescue mission is automatically set; for example, O1 = [0, 2, 5, 1002, 9, 6, 0], and the maximum demand obtained according to the trapped population quantity range at rescue point 2 and rescue point 5 is 50 and 10 respectively. Then, the amount of supplies carried by UAV 1 when departing from the base is 60; the supply demands for rescue point 9 and rescue point 6 according to the trapped population quantity are 30 and 20 respectively. Therefore, the amount of supplies that UAV 1 needs to replenish at supply station 1002 is 50.

[0151] 2.2. Problem Feature-Guided Initialization Strategy

[0152] As mentioned above, the variables that need to be optimized in the present invention are mainly: the task assignment scheme of the UAV, the task sorting, and the selection of the supply time, which correspond to the first row and the second row of the particle respectively. Therefore, the proposed algorithm mainly includes three steps when generating the initialization scheme: (1) Assign all tasks to the UAVs; (2) Sort the tasks assigned to each UAV; (3) Select the time when supplies are needed during the execution of the tasks. For the above steps, specific strategies are given in this section.

[0153] 2.2.1. Initialization Task Assignment Based on Clustering

[0154] For the UAV task assignment problem, the flight distance is an important evaluation index. Therefore, when the UAV continuously executes multiple tasks with similar distances, the flight distance can be shortened, thereby obtaining better results. For this purpose, the present invention proposes an initialization task assignment strategy based on K-means for selecting the values of the first row of the particle.

[0155] First, according to the position information of N rescue tasks, use the K-means clustering algorithm to divide these N tasks into I sub-clusters; then, randomly assign the I sub-clusters to I UAVs to complete the initial task assignment. In particular, to prevent the same clustering result from being assigned multiple times, resulting in serious homogenization of the particles, the present invention sets up an external archive set Ar1 to save different clustering results. In addition, the scale of Ar1 is limited by setting the number of runs of the clustering algorithm, and the general value is the PSO population scale. During the initialization assignment, the clustering results in Ar1 are called in a loop. Each time a loop is executed, the clustering schemes with different results in Ar1 are assigned to the particles in the population that have not been assigned tasks one by one, to avoid the homogenization of the particles to the greatest extent. The particle initialization schematic diagram is as Figure 4 shown.

[0156] 2.2.2 Task sorting based on penalty or time trend

[0157] After completing the initial task assignment in Section 2.2.1, the next step is to sort the rescue tasks of each drone in the particle. This section presents an evaluation metric that takes into account both rescue value (f1) and distance (f2) to calculate the comprehensive benefits of different tasks. Assuming that rescue task j is one of the rescue tasks that drone i is responsible for, the comprehensive benefit of selecting rescue task j as the next rescue task performed by drone i is:

[0158]

[0159] in, represents the comprehensive benefit of drone i choosing rescue mission j as the next rescue mission at time t; v j (t) represents the rescue value of rescue mission j; represents the time it takes for drone i to arrive at rescue mission j from the base; d i,j represents the distance between rescue mission j and drone i; α represents the weight.

[0160] 2.2.3. Probabilistic replenishment strategy based on load

[0161] After completing the task allocation and sorting operations, it is necessary to choose the right time to resupply the drone under the current task execution order. On the one hand, it can prevent the drone from being unable to complete the assigned task due to cargo capacity restrictions. On the other hand, by increasing a certain distance in exchange for increased speed, it can reach the rescue point faster.

[0162] According to the relationship between the drone load and speed in formula (5), the greater the load, the slower the drone’s flight speed. When the drone is empty, its flight speed is the highest, denoted as vel max ; When fully loaded, the UAV’s flight speed is the minimum, recorded as (1-ε)×vel max , where ε is the influence coefficient, which is determined by the performance of the drone itself. Assume that after drone i completes the jth assigned task, the total amount of materials it has unloaded is At this time, the supply probability of UAV i can be expressed as shown in formula (16):

[0163]

[0164] in, represents the probability of resupply of UAV i after completing the jth rescue mission; S max represents the maximum cargo capacity of the drone; in particular, when and When it is the case, it means that UAV i cannot complete the assigned j-th rescue mission. In this situation, after completing the (j - 1)-th rescue mission, the UAV must be resupplied.

[0165] 2.3. Main population evolution

[0166] After the algorithm initialization is completed, perform non-dominated sorting in the form of interval numbers on the initial population Pop, and thus obtain a new population Pop' with a new particle sorting. Then, divide the new population Pop' into the main population Pop1 and the auxiliary population Pop2 according to the ratio θ. Specifically: First, determine the smallest integer θ that can be divided according to the ratio θ based on the principle of the least common multiple int , and then, according to the non-dominated sorting result, select the first θ int particles from the new population Pop' to form a minimum cycle unit, and distribute the particles to Pop1 and Pop2 within the cycle unit according to the ratio θ. After that, continue to select the next θ int particles in sequence for division, and repeat this process until all the particles in Pop' are allocated. For example, in Figure 5 , θ = 2 / 1, θ int = 3, and cycle according to θ int to distribute the individuals in Pop to Pop1 and Pop2 according to the ratio.

[0167] After the population division, the main population and the auxiliary population adopt a simultaneous update strategy. In the main population Pop1, different from the traditional PSO update strategy, in order to avoid over-concentration and homogenization of the population, the setting of the global extreme value is cancelled. The algorithm selects a learning object through the introduction of the auxiliary population Pop2, that is, the learning library, for speed update. The specific update method adopts the method in the literature "Refinement of Collaborative Multitasking and Multi-Objective Allocation in Heterogeneous UAV Swarms An Approach Informed by Decomposition Learning and Particle Swarm Optimization" (hereinafter referred to as literature 2).

[0168] 2.3.1. Selection of learning object

[0169] The present invention selects a learning object through the following steps: First, determine the particles in the learning library that are not dominated by the particle to be updated as the range of the learning object; then, select the particle with the shortest perpendicular distance between the line connecting the particle to be updated and the origin within the range as the learning object. The specific selection example is as Figure 6 shown.

[0170] Figure 6 In (a), the solid circles represent particles. The red color indicates the objects to be updated in the main population, and the green color indicates the range of learning objects that conform to the dominance relationship in the learning library. The numbers below the particles are their dominance levels. Through the straight line L connecting the origin and the particle to be updated, the perpendicular distance between each green particle and L is calculated respectively, and the particle with the shortest distance is selected as the learning object. Therefore, Figure 6 in (a), the green particle marked with a yellow star is the learning object of the red particle. This method can avoid the over-concentration of the population on a single particle, thus preventing the algorithm from falling into a local optimal solution. Figure 6 (b) shows an example of selecting learning objects using this method. It can be seen that not all particles on non-frontiers select particles on the frontier as learning objects, thus ensuring the diversity of the population in iterative updates.

[0171] In addition, using the above method to select learning objects helps high-quality particles have greater advantages when being selected as learning objects, thus accelerating the convergence speed of the algorithm. This can be illustrated by the following proof:

[0172] First, define the similarity: The absolute value of the difference between the f1 / f2 values of two particles A and B is the similarity, as shown in Equation (17):

[0173]

[0174] Assume that X1 is to be updated, X2 and X3 are candidate learning objects, and X2 dominates X3. According to Figure 7 , let the similarities between X2 and X3 and X1 in the objective space be the same, that is, ∠X2Of1 = ∠X3Of1, then △X2OA∽△X3OB. According to the properties of similar triangles, we can get |OX2| < |OX3|, and thus |AX2| < |BX3|. According to the selection criteria of learning objects, X1 is more likely to select particle X2 as the learning object. In addition, since |AX2| = |BC|, when X3 is in the triangular region CX2X3, that is, X3 is more similar to X1 than X2, but |BX3| > |AX2| always holds, so X2 has a higher priority. This type of selection mechanism makes it easier for high-quality particles to be selected as learning objects, thus accelerating the convergence of the algorithm while ensuring diversity.

[0175] 2.3.2. Particle learning and update steps

[0176] After the selection of learning objects is completed, the update process of the particle to be updated will be divided into two stages, and the method proposed in Document 2 is used for iterative update and correction. Figure 8The example illustrates the entire update process: After the particle completes the learning operation, the component UAV2 marked with a green background represents the new allocation plan. However, this plan conflicts with UAV1 and UAV3 (marked with a red background), resulting in the repeated execution of Task 6 and Task 8, and Tasks 1 and 7 are not allocated (marked with a blue background), making the particle in an infeasible position. At the same time, the infeasible particle is corrected (the UAVs other than the learning UAV (UAV2) involved are marked with a yellow background), so that the particle is transferred from the infeasible region to the feasible region, and the update is completed.

[0177] 2.4. Learning Library Evolution Based on Auxiliary Population

[0178] This section introduces the update method of the auxiliary population Pop2 and its auxiliary role for the main population Pop1. The auxiliary population Pop2 serves as a learning library, self-updates through learning, mutation, and screening, and interacts with the main population Pop1.

[0179] 2.4.1. Self-Update Method Inside the Learning Library

[0180] The update of individuals inside the learning library adopts the learning strategy in Section 2.3 as the update operator Δ1; and the genetic mutation operator Δ2 proposed in Document 2 is introduced. During the update, the auxiliary population Pop2 is divided into two blocks, and different update operators are applied to each block. The range of the block is determined randomly, and the specific division is based on the non-dominated rank determined by the random number rank scope The calculation method of the random number is shown in Equation (18).

[0181] rank scope =[1, rand([1, rank max / 2])] (18)

[0182] Among them, rank scope represents the non-dominated rank range of the selected population; rank max is the maximum non-dominated rank of the current generation; rand([1, rank max / 2]) is a random number between [1, rank max / 2]. High-quality individuals within the selected range (non-dominated rank less than rank scope ) are updated in position using the genetic mutation operator Δ2, which helps the algorithm jump out of local optima and search for more solutions. Inferior individuals (non-dominated rank greater than rank scope ) are updated through the update operator Δ1, which helps accelerate the convergence of the population.

[0183] When using the genetic variation operator Δ2 for variation, the guiding direction needs to be selected according to the individual's position. In the present invention, the individual space is divided into two parts, and the individuals are sorted according to the f1 / f2 value. The first half of the individuals are mutated with the orientation of optimizing f1, and the second half are mutated with the orientation of optimizing f2. During the mutation process, the task sorting strategy based on penalty (f1) or time (f2) tendency proposed in Section 2.2.2 is adopted, and the goal is achieved by adjusting the weight parameter α. Figure 9 Fig. Figure 9 shows the evolution process: the blue and black individuals are the individuals to be updated. Among them, the blue individuals are determined to be updated using the Δ2 operator through formula (18), and the blue arrows represent their mutation directions. The green individuals are high-quality individuals that can replace the original individuals after mutation, and the red individuals are low-quality individuals that cannot replace the original individuals. The light blue area represents the area oriented towards optimizing f1, and the light yellow area represents the area oriented towards optimizing f2. The black individuals are updated using the update operator Δ1, and the black arrows represent the learning objects of the individuals to be updated. This method enhances the extension of the Pareto front while ensuring convergence.

[0184] 2.4.2 Interaction between the learning library and the main population

[0185] Since the scale of the learning library is limited and new individuals are intermittently absorbed from the main population during the algorithm iteration update process, an elimination mechanism needs to be introduced to keep the scale of the learning library unchanged.

[0186] According to the content of Section 2.3, when the individuals in the main population select learning objects, they will determine the learning objects according to the non-dominated sorting result of the population inside the learning library and the position of the individual itself. Therefore, non-dominated sorting must be carried out inside the learning library, and the elimination mechanism is also based on the non-dominated sorting result.

[0187] In each iteration update of the learning library, new individuals are generated by the update operator Δ1 or the genetic variation operator Δ2. The new individuals generated by the update operator Δ1 will determine whether to replace the old individuals according to the fitness value, while the new individuals generated by the genetic variation operator Δ2 will jointly participate in non-dominated sorting and crowding degree sorting with the old individuals, and the old individuals with the lowest ranking will be eliminated. The number of eliminated individuals is equal to the number of newly added individuals, so as to keep the scale of the learning library unchanged.

[0188] In addition, to make the most of the learning library, the parameter τ is set to control the exchange frequency between the population inside the learning library and a part of the main population. By default, the value of τ is 1 / 10 of the maximum number of algorithm iterations G.

[0189] 3. Simulation experiments

[0190] 3.1 Evaluation indicators

[0191] The model proposed by the present invention is an uncertain model containing data in the form of interval numbers, where f1 in the objective space is an interval number and f2 is an exact number. During the algorithm iteration process, the numerical processing method proposed in Section 1.4 is adopted, and non-dominated sorting is performed in the form of a mixture of interval numbers and exact numbers. The generated solutions are represented in the objective space as shown in Figure 10 (a). Since the traditional HV value calculation method cannot accurately reflect the quality of the solutions, the present invention uses the Double Hypervolume (DHV) to evaluate the quality of the solutions. According to Figure 10 (a), the DHV value is calculated for the solutions as shown in Figure 10 (b).

[0192] In Figure 10 (a), the red dot is the reference point [f1(r*), f2(r*)]. The particle Xu is composed of both interval numbers and exact numbers, and the corresponding coordinates in the objective space are ([f1(Xu1 - ), f1(Xu1 + )], f2(Xu2)) (the positions of the upper and lower indices need to be adjusted again). Therefore, due to the uncertainty of f1, the HV value cannot be accurately obtained. Only the HV value range of the particle Xu can be determined in region B, while it cannot be determined in regions A and C.

[0193] Therefore, the present invention uses the DHV value shown in Figure 10 (b) for calculation. In the figure, the red dot is the reference point, and the solutions are divided into two categories: the lower-bound solutions composed of blue dots and the upper-bound solutions composed of yellow dots. The HV values of the upper-bound and lower-bound solutions are calculated separately and added together, as shown in formula (19)

[0194]

[0195] Among them, the area of the blue region represents the HV value of the lower-bound solutions, the area of the yellow region represents the HV value of the upper-bound solutions (the yellow region in the figure is completely covered), and the green region is the overlapping part of the two. Using the DHV value to evaluate the quality of the solutions can take into account the influence of the position and uncertainty of the solutions on their quality, and is applicable to the algorithm proposed by the present invention.

[0196] 3.2. Experiments and Parameter Settings

[0197] To verify the superiority of the proposed PILPSO algorithm of the present invention, a comparative experiment was carried out. The specific scenario settings are shown in Table 3. The effectiveness of the key strategy was verified based on the results of 30 experiments. Four advanced multi-UAV task allocation algorithms (AAUC-IBEA, CCPSO, MODABC, and SLPSO) were selected for comparison in the same scenario in the comparative experiment. The DHV value mentioned in Section 3.1 was used as the performance evaluation index in the experiment. The larger the DHV value, the better the algorithm performance. In addition, the Wilcoxon rank-sum test was used to determine whether the PILPSO algorithm was significantly superior to the comparative algorithms, and the confidence level was set to 0.05. Among them, "--" and "++" indicate that the PILPSO algorithm is significantly inferior to and superior to the comparative algorithms respectively, and "≈" indicates that there is no obvious difference between the two.

[0198] Table 3 Scenario Settings

[0199]

[0200] 3. Experimental Results and Analysis

[0201] In this experiment, the proposed PILPSO algorithm was compared with four comparative algorithms: AAUC-IBEA, CCPSO, MODABC, and SLPSO in three scenarios respectively. Table 4 summarizes the results of the five algorithms in the three scenarios.

[0202] Table 4 DHV Values of Comparative Experiments

[0203]

[0204] From the results of Scenario 1, it can be seen that the performance of the proposed PILPSO algorithm of the present invention does not have an advantage compared with MODABC and SLPSO in simple scenarios; however, as the scenario complexity increases, such as in Scenarios 2 and 3, the performance advantage of the PILPSO algorithm gradually appears, and it is completely superior to the other four comparative algorithms in these two scenarios, proving that the PILPSO algorithm has strong optimization ability.

[0205] Further analysis shows that Figure 11 presents the box plots of the DHV values of the proposed method of the present invention and the four algorithms in three scenarios.

[0206] As shown in Figure 11(a), in Scenario 1, the PILPSO algorithm performs better, second only to the MODABC algorithm.

[0207] As shown in Figure 11(b), in Scenario 2, the PILPSO performs the best, and the advantage is relatively obvious compared with the comparative methods.

[0208] As shown in Figure 11(c), in Scenario 3, the scenario is more complex, and the advantage of the PILPSO algorithm is more obvious.

[0209] Figure 11 further verifies that as the scene complexity increases, compared with the other four comparison algorithms, the PILPSO algorithm shows more obvious advantages in performance. The data in Table 4 and Figure 11 both prove that the algorithm proposed in the present invention performs better in the face of complex scenes.

[0210] To further study the algorithm performance, under three scenarios, the Pareto front distribution of the closest experiment to the mean value in 30 experiments of five algorithms is analyzed.

[0211] As shown in Figure 12(a), in Scenario 1, due to the low scene complexity and small feasible region, the Pareto front distributions of the five algorithms have a high degree of overlap. PILPSO performs excellently in distribution and convergence. MODABC is superior to PILPSO in convergence but slightly inferior in distribution. The other three comparison algorithms are inferior to PILPSO in both convergence and distribution.

[0212] As shown in Figure 12(b), in Scenario 2, as the scene complexity increases, the performance advantage of PILPSO gradually emerges. It is completely superior to the other four comparison algorithms in convergence. In terms of distribution, except for no significant difference with MODABC, PILPSO is superior to the other three comparison algorithms.

[0213] As shown in Figure 12(c), in Scenario 3 with the highest scene complexity, the performance gap between PILPSO and the other four comparison algorithms further widens. In terms of convergence, PILPSO shows an absolute advantage. In terms of distribution, although MODABC has a slight advantage in the distribution range, the other algorithms are inferior to PILPSO in both the distribution range and distribution uniformity. At the same time, by observing the width of the particles on the Pareto front of the five algorithms (indicating the uncertainty of the scheme), it can be seen that the uncertainty of the schemes of PILPSO and MODABC is small, while the uncertainty of the other three comparison algorithms is large.

[0214] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0215] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined in the present invention can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in the present invention, but will conform to the widest scope consistent with the principles and novel features disclosed in the present invention.

Claims

1. A method for task allocation of UAV swarm cooperative rescue based on population interaction learning particle swarm algorithm, characterized in that It includes the following steps: Determine the minimum rescue comprehensive penalty objective function based on the rescue value of each rescue mission and the time for each drone to reach the rescue mission it is responsible for from the base; Determine the shortest rescue time objective function based on the time for each drone to reach the last rescue mission it is responsible for; Construct a cooperative rescue mission allocation model for the drone swarm according to the minimum rescue comprehensive penalty objective function and the shortest rescue time objective function; Solve the cooperative rescue mission allocation model for the drone swarm based on the population interactive learning particle swarm algorithm, and allocate the cooperative rescue mission for the drone swarm based on the solution result; The encoding and decoding strategies of the population interactive learning particle swarm algorithm specifically include: In the encoding stage: Integer encoding is adopted, and the particle structure is a two-stage representation method; The first stage is a nested array, which is a matrix of size 1*I, indicating that there are I inner arrays. The inner size of the i-th column array X 1i is 1*n i , representing the rescue tasks assigned to drone i and their execution order. n i is the number of rescue tasks assigned to drone i; assuming there are J rescue tasks in total, so the value range of X 1i is [1, J]; the second stage is an array X i of size 1*ns 2i , representing the supply time of drone i, expressed by the rescue task label, that is, the drone will be resupplied after completing this task. ns i is the number of times of resupply required during the mission of drone i, and its value range is {1, 2,..., n i - 1}; In the decoding stage: A two-stage embedding method is adopted. That is, in the first stage, the starting point, replenishment point, and ending point are embedded according to the array content in the second stage. If decoding the i-th particle X i , first decode the j-th column in the first stage of X i to separate the task assignment of the i-th UAV, that is, a series of task points, denoted as X i ; According to the replenishment point location information given at the corresponding position in the second stage, embed the replenishment point into the task; The embedding operation follows the principle of proximity, that is, select the replenishment point closest to the first rescue task to be executed after the UAV is replenished. Finally, embed the base at both ends of the UAV task assignment sequence to obtain the true flight path of the decoded UAV. This flight path includes the starting point, rescue points, replenishment points, and ending point.

2. The method for allocating collaborative rescue tasks of an unmanned aerial vehicle cluster based on a population interaction learning-based particle swarm optimization algorithm according to claim 1, wherein, The minimum rescue comprehensive penalty objective function is expressed as: Among them, f1 represents the minimum rescue comprehensive penalty objective function; I represents a total of I unmanned aerial vehicles; i represents the i-th unmanned aerial vehicle; J represents a total of J rescue tasks; j represents the j-th rescue task; represents the time for the i-th unmanned aerial vehicle to reach the j-th rescue task from the base; represents the rescue value of the j-th rescue task at time t.

3. The method for allocating UAV swarm cooperative rescue tasks based on the population interaction learning-based particle swarm algorithm according to claim 2, wherein, The corresponding acquisition steps for the time of each drone to reach the rescue mission it is responsible for from the base include: Based on the maximum no-load flight speed of the drone, the maximum load capacity of the drone, and the amount of rescue supplies of each drone before reaching each rescue mission, obtain the flight speed of each drone between every two rescue missions; It is expressed as: Among them, vel i,j represents the flight speed of UAV i between rescue mission j - 1 and rescue mission j; rescue mission j - 1 is the previous rescue mission of rescue mission j; ε represents the UAV performance impact coefficient; vel max represents the maximum no-load flight speed of the UAV; S max represents the maximum cargo capacity of the UAV; S i,j represents the amount of rescue supplies of UAV i before reaching rescue mission j; Based on the flight speed and distance of each drone between every two rescue missions, obtain the time for each drone to reach the rescue mission it is responsible for from the base, which is expressed as: Among them, represents the time for the drone i to reach the rescue mission j from the base; J i represents the number of rescue missions performed by the drone i from leaving the base until returning to the base, and L(j, j - 1) represents the distance between the rescue mission j - 1 and the rescue mission j; ΔT represents the sum of the unloading time and the replenishment time of the drone i.

4. The method for allocating UAV swarm cooperative rescue tasks based on the population interaction learning-based particle swarm algorithm according to claim 1, wherein The shortest rescue time objective function is expressed as: Among them, f2 represents the shortest rescue time objective function; T i end-1 represents the time when the drone i arrives at the last rescue mission it is responsible for; I represents the total number of drones; ΔT represents the sum of the unloading time and replenishment time of the drone i.

5. The method for allocating collaborative rescue tasks for UAV swarms based on the population interaction learning-based particle swarm optimization algorithm according to claim 1, wherein, The solution of the cooperative rescue mission allocation model for the drone swarm based on the population interactive learning particle swarm algorithm specifically includes: Initialize the population interactive learning particle swarm algorithm based on K-means clustering to obtain an initial population; Divide the initial population into a main population and an auxiliary population; Determine the particles to be updated in the main population, and select particle learning objects based on the auxiliary population; Update the particles to be updated based on the particle learning objects.

6. The method for allocating the collaborative rescue tasks of the UAV swarm based on the population interaction learning-based particle swarm algorithm according to claim 5, wherein, The initialization of the population interactive learning particle swarm algorithm based on K-means clustering to obtain an initial population specifically includes: (1) According to the position information of J rescue missions, use the K-means clustering algorithm to divide these J rescue missions into I sub-clusters; randomly assign the I sub-clusters to I drones to complete the initial task allocation; (2) Sort the rescue missions of each drone based on the comprehensive benefit corresponding to the rescue mission responsible for by each drone; Specifically: Let rescue mission j be one of the rescue missions responsible for by drone i, and the comprehensive benefit of selecting rescue mission j as the next rescue mission to be executed by drone i is: Among them, represents the comprehensive benefit of UAV i selecting rescue mission j as the next rescue mission at time t; v j (t) represents the rescue value of rescue mission j; represents the time for UAV i to reach rescue mission j from the base; d i,j represents the distance between rescue mission j and UAV i; α represents the weight; (3) Under the current rescue mission execution order of the drone i, replenish the drone; specifically: assume that after the drone i has completed the jth rescue mission assigned, the total amount of rescue supplies unloaded by the drone i is At this time, the replenishment probability of the drone i is expressed as: Among them, Sr i j represents the resupply probability of the drone i after completing the j-th rescue mission; S max represents the maximum payload of the drone; when Sr i j > 1 and Sr i j-1 < 1, it means that the drone i cannot complete the assigned j-th rescue mission. In this case, after completing the (j - 1)-th rescue mission, the drone must be resupplied.

7. The method for task allocation of UAV swarm cooperative rescue based on population interaction learning particle swarm algorithm according to claim 5, wherein The division of the initial population into a main population and an auxiliary population specifically includes: Perform non-dominated sorting in the form of interval numbers on the initial population Pop to obtain a new population Pop' with a new particle sorting; Determine the smallest integer θ divided according to the ratio θ according to the principle of the least common multiple of integers int ; According to the non-dominated sorting results, select the top θ int particles from the new population Pop' to form a minimum cycle unit, and distribute the particles to the main population Pop1 and the auxiliary population Pop2 within the minimum cycle unit according to the ratio θ; continue to select the next θ int particles for partitioning according to the non-dominated sorting results, and repeat this process until all the particles in the new population Pop' are distributed completely.

8. The method for allocating UAV swarm cooperative rescue tasks based on the population interaction learning-based particle swarm optimization algorithm according to claim 5, wherein, Determine particle learning objects based on the main population and the auxiliary population, specifically including: Take the particles in the auxiliary population that are not dominated by the particles to be updated as the range of particle learning objects; Select the particle with the shortest perpendicular distance between the line connecting the particle to be updated and the origin as the particle learning object within the range.

Citation Information

Patent Citations

  • Heterogeneous unmanned aerial vehicle cluster multi-task allocation method based on decomposition learning particle swarm

    CN118466582A