A multi-uav task allocation method and system for offshore oil film monitoring
By optimizing UAV task allocation through a hybrid learning pigeon flocking optimization algorithm and Q-learning mechanism, the problem of collaborative task allocation of UAVs in dynamic environments in marine oil spill monitoring was solved. This enabled the synchronous arrival of fixed-wing and multi-rotor UAVs and the dynamic redistribution of oil film drift, thereby improving monitoring accuracy and the robustness of the scheme.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
- Filing Date
- 2026-06-22
- Publication Date
- 2026-07-21
AI Technical Summary
Existing multi-UAV task allocation methods are difficult to adapt to dynamic environments in marine oil spill monitoring, resulting in inconsistent oil film thickness inversion accuracy, and the collaborative task allocation of fixed-wing and multi-rotor UAVs is difficult to meet time constraints.
A hybrid learning pigeon flock optimization algorithm is adopted, which combines Q-learning mechanism and priority bidding scheduling allocation mechanism to construct a multi-UAV task allocation model. By dynamically sensing population diversity and temporal satisfaction rate through landmark operator, the UAV task allocation scheme is optimized, and a collaborative temporal repair operator is introduced to ensure the arrival time synchronization of fixed-wing and multi-rotor UAVs.
It achieves rapid and robust UAV task allocation in dynamic environments, improves the spatiotemporal coverage accuracy and data consistency of oil film monitoring, adapts to the dynamic redistribution capability of oil film drift, and ensures the feasibility of task allocation schemes and the synchronization of UAV arrival times.
Smart Images

Figure CN122434073A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to multi-UAV task allocation, specifically to a multi-UAV task allocation method and system for marine oil slick monitoring. Background Technology
[0002] Marine oil spills pose a serious threat to the marine ecosystem, and rapid and accurate monitoring of the spill's extent and thickness is crucial for emergency response. Unmanned aerial vehicles (UAVs) equipped with hyperspectral and thermal infrared sensors have become an important means of oil spill monitoring due to their maneuverability and low cost. However, a single UAV is limited by its field of view and endurance, making it difficult to independently complete accurate measurements of large-scale oil spills. Multi-UAV collaborative monitoring can effectively improve coverage efficiency and data accuracy. Fixed-wing UAVs are fast and have long flight times, making them suitable for large-area scanning; multi-rotor UAVs can hover and perform precise measurements, making them suitable for fixed-point thickness inversion. Achieving heterogeneous collaboration between these two types of UAVs can effectively solve this problem.
[0003] Existing research employs Pigeon-Inspired Optimization (PIO) for multi-UAV task allocation, but this assumes that the UAVs are isomorphic and the task points are independent. While Pigeon-Inspired Optimization has advantages such as fast convergence and strong robustness, its landmark operator uses a fixed elimination ratio, making it difficult to adapt to changes in the population state under dynamic environments.
[0004] In addition, marine oil spill monitoring faces the challenge of continuous oil film drift. Existing task allocation methods mostly assume a static environment and lack a dynamic response mechanism to drift, resulting in spatiotemporal inconsistencies in actual monitoring data and affecting the accuracy of oil film thickness inversion. Summary of the Invention
[0005] Purpose of the invention: To address the above-mentioned shortcomings, this invention provides a multi-UAV task allocation method and system for marine oil slick monitoring that adapts to dynamic environments.
[0006] Technical solution: To solve the above problems, the present invention provides a multi-UAV task allocation method for marine oil slick monitoring, comprising the following steps:
[0007] (1) Obtain the extent of the oil slick at sea, preliminarily determine the monitoring points for oil slick monitoring, and use multiple UAVs to monitor the monitoring points;
[0008] (2) With the goal of minimizing the energy consumption of multiple UAVs, and with the UAVs’ endurance, arrival time and oil film drift tolerance as constraints, a multi-UAV task allocation model is constructed.
[0009] (3) Based on the location of monitoring points, UAV parameters and multi-UAV task allocation model of oil film monitoring, the task sequence of UAVs visiting monitoring points is used as the optimization variable. Combined with the objective function, the task allocation is carried out by the hybrid learning pigeon flock optimization algorithm, and the task allocation scheme is output. The hybrid learning pigeon flock optimization algorithm introduces the Q-learning mechanism in the second stage of the pigeon flock algorithm. According to the population diversity and temporal satisfaction rate, the elimination ratio and elimination individual type are determined. The landmark operator is updated according to the surviving individuals after elimination. The pigeon flock algorithm is carried out based on the updated landmark operator.
[0010] Furthermore, the method of using multiple drones to monitor a monitoring point includes: using one fixed-wing drone and one multi-rotor drone to monitor a monitoring point, wherein the arrival time of the fixed-wing drone at the monitoring point is earlier than the arrival time of the multi-rotor drone, and the arrival time difference between the fixed-wing drone and the multi-rotor drone at the same monitoring point is within a preset threshold.
[0011] Furthermore, the multi-UAV task allocation model is as follows:
[0012] ;
[0013] in, Let be the objective function. Total flight distance Penalties for timing violations of arrival times for fixed-wing and multi-rotor drones. For load balancing variance, , , These are the weighting coefficients. Indicates the first monitoring points Assigned to the fixed-wing drones Assignment variables, For the number of fixed-wing drones, Indicates the first monitoring points Assigned to the Multi-rotor drones Assignment variables, For the number of monitoring points, For the number of multi-rotor drones, Represents the set of monitoring points. This indicates the locations of all drone take-off and landing bases. Indicates monitoring point The initial geographical coordinates, Indicates fixed-wing unmanned aerial vehicle The first monitoring point visited, Represents Euclidean distance. Indicates monitoring point The initial geographical coordinates, Indicates fixed-wing unmanned aerial vehicle The first visit One monitoring point, Indicates monitoring point The initial geographical coordinates, Indicates fixed-wing unmanned aerial vehicle The first visit One monitoring point, Indicates fixed-wing unmanned aerial vehicle The total number of monitoring points allocated, Indicates fixed-wing unmanned aerial vehicle Maximum range Indicates monitoring point The initial geographical coordinates, Indicates multi-rotor drone The first monitoring point visited, Indicates monitoring point The initial geographical coordinates, Indicates multi-rotor drone The first visit One monitoring point, Indicates monitoring point The initial geographical coordinates, Indicates multi-rotor drone The first visit One monitoring point, Indicates multi-rotor drone The total number of monitoring points allocated, Indicates multi-rotor drone Maximum range This indicates that you have been assigned to a monitoring point. Multi-rotor drones arrive at monitoring points At that moment, This indicates that you have been assigned to a monitoring point. The fixed-wing drone arrived at the monitoring point At that moment, The maximum allowable time difference, Indicates the oil film drift tolerance. Indicates the oil film drift velocity. Indicates monitoring point exist Geographic coordinates at any given time Indicates monitoring point exist Geographic coordinates at any given time Indicates monitoring point exist Geographic coordinates at any given time This indicates the maximum number of monitoring points that each fixed-wing drone can visit in a single launch. This indicates the maximum number of monitoring points that each multi-rotor drone can visit in a single launch. Indicates fixed-wing unmanned aerial vehicle Arrival at the monitoring point At that moment, Indicates multi-rotor drone Arrival at the monitoring point At that moment.
[0014] The total flight distance The calculation formula is:
[0015] ;
[0016] in, Indicates fixed-wing unmanned aerial vehicle The geographical coordinates of the first monitoring point visited. Indicates fixed-wing unmanned aerial vehicle The first visit The geographical coordinates of each monitoring point Indicates fixed-wing unmanned aerial vehicle The first visit The geographical coordinates of each monitoring point Indicates multi-rotor drone The geographical coordinates of the first monitoring point visited. Indicates multi-rotor drone The first visit The geographical coordinates of each monitoring point Indicates multi-rotor drone The first visit The geographical coordinates of each monitoring point;
[0017] The timing violation penalty for the arrival time of the fixed-wing UAV and multi-rotor UAV The calculation formula is:
[0018] ;
[0019] The load balancing variance The calculation formula is:
[0020] ;
[0021] in, For fixed-wing unmanned aerial vehicles Total flight distance This represents the average flight distance of all fixed-wing UAVs. For multi-rotor drones Total flight distance This represents the average flight distance of all multi-rotor drones.
[0022] Furthermore, the step of task allocation using the hybrid learning pigeon flock optimization algorithm includes:
[0023] Obtain the initial position and initial speed of the pigeons, and determine the initial population based on the initial position and initial speed;
[0024] The fitness value of each pigeon in the initial population is calculated using the objective function, and the position of the target pigeon is determined based on the fitness value of each pigeon. The population diversity and temporal constraint satisfaction rate are also calculated.
[0025] Based on the location of the target pigeon, the position and speed of each pigeon are updated using a geomagnetic operator;
[0026] After the number of iterations reaches the first value, the current state is determined based on population diversity and temporal constraint satisfaction rate. An action is selected based on the Q value. The action includes elimination ratio and elimination individual type. The elimination ratio determines the individual with the worst fitness in the current population to be eliminated. The elimination individual type determines whether to eliminate directly or replace the eliminated individual. Direct elimination means directly deleting the eliminated individual and randomly copying it from the surviving individuals to supplement the original population size. Replacement elimination means randomly selecting two parents from the surviving individuals, performing single-point crossover between the two parents to generate two offspring, and randomly selecting one of the two offspring to replace the individual to be eliminated. The action includes: directly eliminating or replacing the determined individual to be eliminated.
[0027] The landmark operator is updated based on the center position of the surviving individuals after the action is performed; the position of each pigeon is updated based on the updated landmark operator; the fitness value of each pigeon is calculated; and the global optimal fitness is updated.
[0028] After performing an action, observe the new state and calculate the reward, then update the Q value;
[0029] If the iteration stopping condition is met, output the position of the pigeon with the global optimal fitness, and determine the task allocation scheme based on the pigeon position.
[0030] Furthermore, the position of each pigeon is updated based on the updated landmark operator as follows:
[0031] ;
[0032] in, for The position vector of the pigeon at any given moment. for The position vector of the pigeon at any given moment. The numbers are uniformly random. The center position of the surviving individual after the action was performed:
[0033] ;
[0034] in, Indicates surviving individuals, The maximum fitness value among surviving individuals. For the first The position vectors of the pigeons. For the first The fitness value of each pigeon.
[0035] Furthermore, an auction mechanism is used to convert pigeon locations into task allocation schemes. The steps for determining the task allocation scheme based on pigeon locations are as follows:
[0036] Position the pigeon in the middle front Each component corresponds to a bid price for a fixed-wing UAV; later Each component corresponds to a bid price for a multi-rotor drone;
[0037] All fixed-wing UAVs are sorted from highest to lowest bid price to obtain the first ordered list; all multi-rotor UAVs are sorted from highest to lowest bid price to obtain the second ordered list; and each monitoring point is sorted from highest to lowest oil film thickness or urgency to obtain the third ordered list.
[0038] Based on the sorting of the third ordered list, combined with the multi-UAV task allocation model, fixed-wing UAVs and multi-rotor UAVs are selected from the first ordered list and the second ordered list for each monitoring point in turn; and the remaining endurance and the number of assigned tasks of each fixed-wing UAV and multi-rotor UAV are updated.
[0039] For each drone, its assigned monitoring points are arranged from closest to furthest from the base to obtain the access order, which is the task allocation scheme.
[0040] Furthermore, the task allocation scheme is partially modified to ensure that the arrival time difference between fixed-wing UAVs and multi-rotor UAVs at each monitoring point meets a preset constraint; the partial modification actions include:
[0041] Action 1: Insert a virtual detour point between the target monitoring point and the previous monitoring point in the drone's access sequence. The virtual detour point is located on the perpendicular bisector of the line connecting the previous monitoring point and the target monitoring point, thereby increasing the drone's flight distance, and ensuring that the total flight distance after detour does not exceed its remaining endurance.
[0042] Action 2: Move the target monitoring point forward in the drone's access sequence, or remove unnecessary detour points in the drone's path;
[0043] Action 3: Transfer the monitoring task of the target monitoring point to other drones that are idle or have a task load less than the threshold;
[0044] The partial repair process is as follows:
[0045] If the arrival order of the fixed-wing drone and the multi-rotor drone is reversed, the fixed-wing drone will perform action 2. If this is ineffective, the multi-rotor drone will perform action 1. If this is ineffective, either the fixed-wing drone or the multi-rotor drone will perform action 3.
[0046] If the arrival time difference between the fixed-wing drone and the multi-rotor drone is greater than the maximum allowable time difference, the multi-rotor drone will perform action 2. If this is ineffective, the fixed-wing drone will perform action 1. If this is also ineffective, either the fixed-wing drone or the multi-rotor drone will perform action 3.
[0047] If the arrival time difference between the fixed-wing UAV and the multi-rotor UAV meets the preset constraints after the repair, the repaired access order will be used as the new task allocation scheme; otherwise, the original task allocation scheme will be retained.
[0048] Furthermore, it also includes step (4), at each interval After a few seconds, the geographical coordinates of each monitoring point that the drone has not reached are determined again. If the difference between the geographical coordinates of the monitoring point at this time and the geographical coordinates at the previous time is greater than the tolerance threshold, the monitoring point is marked as invalid and the invalid monitoring point in the drone access sequence is released.
[0049] Step (3) is repeated for the failed monitoring points and idle drones to obtain the redistribution results. The redistribution results are then merged with the original task allocation scheme to form a new task allocation scheme. After each redistribution, the oil film drift speed is updated.
[0050] The present invention discloses a multi-UAV task allocation system for marine oil slick monitoring, comprising:
[0051] The data acquisition module is used to obtain the extent of the oil slick at sea, preliminarily determine the monitoring points for oil slick monitoring, and use multiple UAVs to monitor the monitoring points;
[0052] The model building module is used to construct a multi-UAV task allocation model with the objective function of minimizing the energy consumption of multiple UAVs and the constraints of UAV endurance, arrival time and oil film drift tolerance.
[0053] The task allocation module is used to allocate tasks based on the location of monitoring points, UAV parameters, and a multi-UAV task allocation model for oil film monitoring. It takes the task sequence of UAVs visiting monitoring points as an optimization variable, combines it with the objective function, and uses a hybrid learning pigeon flocking optimization algorithm to allocate tasks and output the task allocation scheme. The hybrid learning pigeon flocking optimization algorithm introduces a Q-learning mechanism in the second stage of the pigeon flocking algorithm. Based on the population diversity and temporal satisfaction rate, it determines the elimination ratio and the type of eliminated individuals. It updates the landmark operator based on the surviving individuals after elimination and performs the pigeon flocking algorithm based on the updated landmark operator.
[0054] Beneficial effects: Compared with the prior art, the significant advantages of this invention are:
[0055] By employing the Q-learning mechanism, the landmark operator dynamically perceives population diversity and temporal satisfaction rate, and makes online decisions on the elimination ratio and elimination type. This achieves an autonomous balance between algorithm exploration and development, improves the search efficiency and convergence quality during the local fine-tuning stage, and adapts to changes in population state under dynamic environments.
[0056] The priority bidding scheduling and allocation mechanism and the collaborative timing repair operator are introduced into the pigeon flock optimization, replacing the single position update method of the traditional pigeon flock algorithm. This ensures the feasibility of the task allocation scheme and the synchronization of arrival time between fixed-wing and multi-rotor aircraft, and avoids the scheme from violating hard timing constraints.
[0057] This invention can quickly and robustly solve the problem of spatiotemporal collaborative task allocation for heterogeneous UAVs in marine oil spill monitoring, and also has the ability to dynamically redistribute tasks based on oil film drift, providing a new hybrid learning optimization approach for multi-UAV collaborative task allocation in complex dynamic environments. Attached Figure Description
[0058] Figure 1 This is a flowchart illustrating the task allocation method in this invention. Detailed Implementation
[0059] like Figure 1 As shown in this embodiment, a multi-UAV task allocation method for marine oil slick monitoring includes the following steps: (1) Obtain the extent of the oil slick at sea, preliminarily determine the monitoring points for oil slick monitoring, and use multiple UAVs to monitor the monitoring points; (2) With the goal of minimizing the energy consumption of multiple UAVs, and with the UAVs’ endurance, arrival time and oil film drift tolerance as constraints, a multi-UAV task allocation model is constructed. (3) Based on the location of monitoring points, UAV parameters and multi-UAV task allocation model of oil film monitoring, the task sequence of UAVs visiting monitoring points is used as the optimization variable. Combined with the objective function, the task allocation is carried out by the hybrid learning pigeon flock optimization algorithm, and the task allocation scheme is output.
[0060] Further, specific steps include:
[0061] Step 1: Obtain the extent of the oil slick at sea, preliminarily determine the monitoring points for oil slick monitoring, and use multiple UAVs to monitor the monitoring points.
[0062] For marine oil spill monitoring, a heterogeneous fleet of fixed-wing and multi-rotor UAVs is used. In this embodiment, each monitoring point requires one fixed-wing UAV and one multi-rotor UAV to work together: the fixed-wing UAV arrives quickly first and performs coarse positioning, while the multi-rotor UAV arrives later to perform hovering and fine measurement. The arrival time difference between the two must be strictly controlled within a preset threshold to ensure that oil film drift does not affect the accuracy of data fusion. The location of the monitoring point is known (obtained through pre-scanning), but it may drift slowly due to the influence of sea surface winds and currents.
[0063] Step 2: Construct a multi-UAV task allocation model with the objective function of minimizing the energy consumption of multiple UAVs and the constraints of UAV endurance, arrival interval time, and oil film drift tolerance.
[0064] Step 2.1, determine the basic definitions for drone task allocation, specifically:
[0065] (1) Let the set of fixed-wing UAVs be:
[0066]
[0067] In the formula, The number of fixed-wing UAVs is given, and their flight speed is constant. (m / s), maximum range is (m).
[0068] (2) Let the set of multi-rotor UAVs be:
[0069]
[0070] In the formula, The number of multi-rotor drones is given, and their flight speed is constant. (m / s), maximum range is (m), and .
[0071] (3) Assume that all UAVs have the same take-off and landing base, denoted as Its coordinates are known. Let the oil film drift velocity vector be... Within a short time window, it can be approximated as a constant vector. Let the set of monitoring points be:
[0072]
[0073] In the formula, The number of monitoring points is given, and the initial geographical coordinates of each monitoring point are known, denoted as:
[0074]
[0075] In the formula, The x-coordinate of the monitoring point The vertical coordinate of the monitoring point is denoted as y.
[0076] Step 2.2, determine the decision variables, specifically:
[0077] (1) Assignment variables for monitoring points and UAVs:
[0078]
[0079] Indicates monitoring point Should it be assigned to a fixed-wing drone? .
[0080]
[0081] Indicates monitoring point Should it be assigned to a multi-rotor drone? .
[0082] (2) Task sequence variables for each UAV:
[0083] For fixed-wing UAVs The set of monitoring points assigned to it is as follows:
[0084]
[0085] Assume the access order is as follows: .
[0086] multi-rotor drones The set of monitoring points assigned to it is as follows:
[0087]
[0088] Assume the access order is as follows: . , For arrangement.
[0089] (3) Arrival time variable:
[0090] Indicates fixed-wing unmanned aerial vehicle Arrival at the monitoring point The moment (if) );
[0091] Indicates multi-rotor drone Arrival at the monitoring point The moment (if) ).
[0092] Step 2.3, calculate the time it takes for the drone to arrive at the monitoring point, specifically:
[0093] (1) Arrival time extrapolation of fixed-wing UAVs:
[0094] Departing from the base, fixed-wing drones The time to reach its first mission point is:
[0095] (1)
[0096] In the formula, The distance is Euclidean. The time progression for subsequent mission points is as follows:
[0097] (2)
[0098] (2) The arrival time of the multi-rotor UAV is recursively calculated, and similarly, we have:
[0099] (3)
[0100] (4).
[0101] Step 2.4, determine the objective function for multi-UAV cooperative task allocation, specifically:
[0102] The optimization objective is to minimize the weighted multi-objective factor.
[0103] (5)
[0104] In the formula, Total flight distance (energy consumption). Penalty for timing violation For load balancing variance.
[0105] (6)
[0106] (7)
[0107] (8)
[0108] In the formula, , monitoring points The allocated fixed-wing and multi-rotor drones, The penalty coefficient for sequential violation, The maximum allowable time difference is determined by the drift speed. and spatial tolerance Confirmed; among them For fixed-wing unmanned aerial vehicles Total flight distance This represents the average flight distance of all fixed-wing drones, and the same applies to multi-rotor drones.
[0109] Let the weighting coefficients satisfy:
[0110] (9).
[0111] Step 2.5, determine the constraints, specifically:
[0112] (10)
[0113] Wherein, equation (a) indicates that each monitoring point needs to be allocated one fixed-wing UAV and one multi-rotor UAV; equation (b) indicates that the total flight distance of each fixed-wing UAV must not exceed its maximum endurance; equation (c) indicates that the total flight distance of each multi-rotor UAV must not exceed its maximum endurance; and equation (d) indicates that for each monitoring point... All must ensure that the fixed-wing UAV arrives before the multi-rotor UAV, and that the arrival time difference does not exceed the oil film drift tolerance, where the maximum time difference is... Drift speed and spatial tolerance Equation (e) indicates that, considering oil film movement, the actual position of the monitoring point changes over time, and ensuring spatial matching of data acquired by both machines, dynamic updates are provided. The physical basis; Equation (f) represents the maximum number of times each drone can perform a single launch. One monitoring point.
[0114] The multi-UAV task allocation model is as follows:
[0115] .
[0116] Step 3: Use a hybrid learning pigeon flock optimization algorithm based on auction-time strategy to allocate tasks.
[0117] Based on the standard pigeon flock optimization algorithm, this invention proposes an Auction-Temporal Hybrid Learning Pigeon-Inspired Optimization (ATHL-PIO) algorithm based on an auction-temporal strategy. This algorithm achieves efficient solution of spatiotemporal collaborative task allocation for heterogeneous UAVs through three core mechanisms: (1) random initialization; (2) dynamic adjustment of elimination ratio and breeding strategy by an adaptive landmark operator based on Q-learning to balance global exploration and local development; (3) real number encoding, using a priority auction scheduling allocation mechanism to transform continuous position vectors into allocation schemes that satisfy hard constraints, so that the allocation schemes satisfy equations (a), (b), (c) and (f) in the constraint conditions; (4) the collaborative temporal optimization operator actively repairs the arrival time difference between fixed-wing UAVs and multi-rotor UAVs, so that the allocation schemes satisfy equation (d) in the constraint conditions.
[0118] Step 3.1, Initialization: Pigeon population size is... Each pigeon represents a candidate solution, and each component of each pigeon is in The internal components are randomly generated according to a uniform distribution, with an initial velocity... Each component in Randomly generated within. (Number) The position vector of the pigeon is denoted as:
[0119] .
[0120] Step 3.2, according to the definition of the objective function... The adaptability of the pigeons:
[0121] (11)
[0122] in, Calculate the formulas and penalty terms in the objective function separately. A smaller fitness value corresponds to a better solution. The algorithm always records the globally optimal fitness. and the corresponding pigeon locations At the same time, population diversity must also be recorded:
[0123] (12)
[0124] In the formula, For fitness standard deviation, The initial population standard deviation is given. The temporal constraint satisfaction rate is also recorded.
[0125] (13).
[0126] Step 3.3: The geomagnetic operator (first stage) performs global coarse allocation. Let the maximum number of iterations in the first stage be... Initial velocity exist Randomly generated from the middle, finally resulting in one dimensional velocity vector , It is the total number of fixed-wing drones and multi-rotor drones, that is, the total dimension of the pigeon's position vector.
[0127] In the updated rules, for each pigeon ( ):
[0128] (14)
[0129] (15)
[0130] In the formula, The geomagnetic attenuation coefficient, These are uniformly random numbers.
[0131] After the update, boundary constraints are applied to each component of the position vector:
[0132] (16)
[0133] In the formula, Indicates the first Only one pigeon in the 1st In the nth iteration, its position vector is the th Each component (a specific bid price corresponding to that pigeon). After each location update, a priority bidding and scheduling allocation mechanism is used to allocate... Transform into a task allocation scheme and calculate fitness. ;like Then update and When the number of iterations reaches At that time, the geomagnetic phase ends.
[0134] The priority bidding and allocation mechanism uses a real number encoding method to directly map the position vector of each pigeon to the "bid price" vector of multiple fixed-wing drones and multi-rotor drones.
[0135] For the The position vector of the pigeon:
[0136]
[0137] In the formula, the first Each component corresponds to a bid price for a fixed-wing UAV; later Each component corresponds to a bid price for a multi-rotor drone. A higher bid price indicates that the drone has a higher priority in the bidding process.
[0138] Map pigeon locations to task assignment schemes. Input pigeon locations. Monitoring point set And the performance parameters of the drone, outputting the allocation variables. , and task sequence , The process of this mechanism proceeds sequentially according to the priority of the monitoring points (e.g., from largest to smallest oil film thickness). The specific steps are as follows:
[0139] (1) Initialization: All fixed-wing UAVs are set at the bid price. Sort from highest to lowest to obtain an ordered list. Similarly, all multi-rotor drones will be priced according to the bid price. Sort from highest to lowest to get .
[0140] (2) Priority ranking of monitoring points: Calculate the oil film thickness or urgency of each monitoring point and arrange them in descending order. .
[0141] (3) Cyclic allocation: For each monitoring point ,from The first unassigned fixed-wing drone in sequence, whose maximum number of mission points has not been exceeded and whose remaining range is sufficient to travel from the base to that point and back (or to the next assignment point), will be selected. If the current drone's remaining battery life is insufficient, proceed to the next drone; this continues until a viable fixed-wing drone is found. The same applies to selecting a multi-rotor drone. Update the drone's remaining battery life and allocated points.
[0142] (4) Sequence generation: For each UAV, arrange its mission points in order of distance from the base from closest to farthest to obtain the access order. , This yields the task allocation scheme. Further, subsequent steps will use timing optimization operators to adjust this order and optimize the task allocation scheme.
[0143] During the allocation process, the battery life and capability constraints are verified in real time, and each monitoring point is forcibly allocated one fixed-wing drone and one multi-rotor drone. If there are no available drones due to the bidding price ranking (an extreme case), the unallocated monitoring point is retained and a severe penalty is imposed on the fitness.
[0144] Step 3.4: Based on the landmark operator (second stage) of the standard pigeon flock optimization algorithm, a tabular Q-learning mechanism is introduced, which enables the algorithm to dynamically determine the elimination ratio in each generation of landmark operators and whether to enable crossbreeding based on the diversity of the current population and the temporal constraint satisfaction rate.
[0145] Let the maximum number of iterations in the second stage be... .
[0146] (1) State space. The state is determined by population diversity. and timing constraint satisfaction rate Discretization.
[0147] Will and Discretize into two levels respectively: The level is low. The level is high; The level is low. The level is high. and If the value is a settable constant, then the total state space size is 4 discrete states, denoted as:
[0148] .
[0149] (2) Action space. In each state The Q-learning controller then selects an action from the following set of actions. Elimination rate and Elimination of Individual Types The elimination type includes direct elimination and replacement elimination. If the total number of actions is 6, then the action space is denoted as:
[0150] .
[0151] Elimination rate : Select the least fit species in the current population Proportional individual removal. Elimination individual type: If true, replacement elimination is used, and two parents are randomly selected from the surviving individuals for single-point crossover to generate two new individuals. One of the two new individuals is randomly selected to replace the eliminated individual; if false, direct elimination is used, the eliminated individual is directly deleted, and a random copy is made from the surviving individuals to supplement the original population size.
[0152] (3) Reward function. Execution action. Receive instant rewards afterwards The reward function comprehensively considers convergence, temporal constraints, and diversity:
[0153] (17)
[0154] In the formula, This represents the improvement amount for the globally optimal fitness (positive values are rewards, negative values are penalties). This represents the improvement in the time constraint satisfaction rate. To minimize the desired population diversity, when Applying penalties ; This means that if the selected action causes the average fitness of the population to decrease instead of increase and the diversity to continue to deteriorate after this iteration, a fixed penalty will be applied; These are the weighting coefficients for each item.
[0155] (4) Q-value initialization and update. Initialize the Q-value, setting all values to 0. The landmark operator is used in each round. -Greedy strategy selects actions based on probability. Randomly select actions, with probability Choose the action that maximizes the Q value in the current state. The exploration rate will be linearly decayed.
[0156] In each iteration, based on the current state Select Action Observe after execution and rewards Update Q value:
[0157] (18)
[0158] In the formula, For learning rate, This is the discount factor.
[0159] (5) Adaptive landmark operator execution steps.
[0160] First, calculate the current state. Based on the current Q value and -greedy strategy selects action Based on the elimination ratio during the exercise, the pigeons are first sorted by fitness level to determine which ones should be eliminated. Proportional individuals; Determine the type of individuals to be eliminated:
[0161] like For each eliminated individual, two parents are randomly selected from the surviving individuals, and two offspring are generated by single-point crossover to replace the eliminated individual.
[0162] like If an individual is eliminated, it is directly deleted, and the surviving individuals are randomly copied to replenish the original population size. ;
[0163] Calculate the weighted center position of surviving individuals (including newly generated individuals). (Same as standard landmark operator):
[0164] (19)
[0165] In the formula, The maximum fitness value (worst value) among the surviving individuals;
[0166] All individuals move toward the center:
[0167] (20)
[0168] In the formula, The result is a uniform random number. After the update, the components of the position vector are subjected to boundary constraints in the same way as in equation (17).
[0169] Based on the updated pigeon positions, a priority auction scheduling allocation mechanism is adopted to convert the task allocation scheme. The fitness is calculated based on the task allocation scheme, and the global optimal fitness is updated.
[0170] Calculate reward after the action is performed Observe the new state Update the Q value;
[0171] Decay Exploration Rate Then proceed to the next iteration. If the iteration stops, output the position of the pigeon with the best fitness in the world. Based on the pigeon position, use a priority bidding scheduling mechanism to determine the task allocation scheme.
[0172] Step 3.5, to meet the timing consistency requirements (i.e., to avoid the arrival time difference between fixed-wing and multi-rotor aircraft). Exceeding (Or the order may be reversed). After each position update and decoding, an operator with local repair capabilities is introduced, called the cooperative timing optimization operator. This operator is used to actively adjust the access order in the UAV task allocation scheme so that the arrival time difference between fixed-wing UAVs and multi-rotor UAVs at each monitoring point meets the preset constraints. .
[0173] (1) For fixed-wing unmanned aerial vehicles and multi-rotor drones Arrive at the monitoring point The estimated times are respectively , Input: Allocation matrix , Initial task sequence , Coordinates of each UAV take-off and landing base Flight speed , and the maximum allowable time difference for oil film drift .
[0174] (2) Repair Action 1: Drone Path Delay. This occurs when the drone arrives at the monitoring point. A previous mission point and Insert a virtual detour point between Its location is and The distance from the original straight line increases along the perpendicular bisector of the connecting line. rice( (determined by binary search), such that Increase the distance until the time difference limit is met, while ensuring that the total flight distance of the drone after the detour does not exceed its remaining range. If a feasible... Then update the drone's sequence and insert the detour point.
[0175] Action 2: Advance the drone's path. Set the monitoring points... Advance the drone's mission sequence to allow it to arrive earlier; or eliminate unnecessary detours in the drone's path to shorten its journey. Flight distance.
[0176] Action 3: Cooperative Pairing and Exchange. If the above local sequence adjustment cannot meet the timing constraints, collect the set of currently idle or unassigned fixed-wing or multi-rotor UAVs, select a UAV whose arrival time meets the timing constraints (and whose own endurance is sufficient after replacement), and transfer the task of the monitoring point. If successful, update the task sequence of the corresponding UAV; otherwise, the repair fails.
[0177] (3) Repair process: If If the arrival order of the fixed-wing drone and the multi-rotor drone is reversed, the fixed-wing drone will perform action 2 first. If this is ineffective, the multi-rotor drone will perform action 1. If this is also ineffective, either the fixed-wing drone or the multi-rotor drone will perform action 3. If the multi-rotor drone is ineffective, then the fixed-wing drone will perform action 1. If the fixed-wing drone is ineffective, then either the multi-rotor drone or the fixed-wing drone will perform action 3.
[0178] (4) If the repair action is successfully found, If the task allocation scheme (sequence or pairing) fails, the task allocation scheme is updated, and the next monitoring point is processed. If all repair actions fail to bring the time difference within the allowable range, the original scheme is retained, and the timing violation cost of that monitoring point will be borne by the fitness function. This is directly reflected in the text.
[0179] Step 4: Add a dynamic diffusion-aware redistribution mechanism to ensure that the allocation scheme satisfies equation (e) in the constraint conditions, and output the optimal task allocation scheme.
[0180] Step 4.1, set the oil film drift velocity. It remains constant in the short term and can be measured using real-time meteorological data. Monitoring points At any moment The predicted location is:
[0181] (twenty one)
[0182] In the formula, The time when the initial plan is completed (usually set) If the speed changes during task execution, a piecewise linear model is used to predict the position:
[0183] (twenty two)
[0184] Step 4.2, the system every... Check once per second for each monitoring point that has not yet been executed. (i.e., the point that the drone has not yet reached), if the following conditions are met:
[0185] (twenty three)
[0186] Then mark the monitoring point as "failed". This is the tolerance threshold. Meanwhile, if the estimated arrival time of a certain monitoring point... The difference from the current time is less than Furthermore, if the drift distance exceeds the limit, the redistribution mechanism can be triggered in advance.
[0187] Step 4.3: When a monitoring point is marked as "failed", perform the following steps:
[0188] (1) Collect all idle drones (including those released after failure and those originally without assigned tasks) to form a temporary drone ensemble.
[0189] (2) Form a new task set by combining the failed monitoring points, and use an auction allocation mechanism to redistribute the temporary UAV set. During redistribution, priority is given to monitoring points with large oil film thickness and fast drift speed.
[0190] (3) Merge the redistribution results with the original allocation scheme of the non-failed points to generate a new global task plan.
[0191] (4) Adjust the sequence by using a collaborative time series optimization operator to meet the time series consistency constraint.
[0192] After each redistribution, the calculation is recalculated based on the current oil film drift velocity. Update to all subsequent calculations. If the drift speed changes drastically, the inspection interval can be shortened. To improve response speed It can learn in Q-learning, and the Q-learning controller will adaptively adjust the checking frequency based on the drift speed and the timing constraint satisfaction rate.
[0193] When the pigeon flock optimization algorithm reaches its maximum number of iterations If the global optimal fitness remains unchanged for dozens of generations, output the task allocation scheme corresponding to the global optimal solution: a list of monitoring points each UAV is responsible for and the order in which they are visited; the estimated time difference between each pair of fixed-wing UAVs and multi-rotor UAVs arriving at the same monitoring point; and the initial check interval required for reassignment. and drift velocity threshold.
[0194] In practice, the dynamic reallocation mechanism will continuously monitor and adjust the plan to ensure the spatiotemporal synchronization accuracy and robustness of the marine oil spill monitoring mission.
Claims
1. A multi-UAV task allocation method for marine oil slick monitoring, characterized in that, Includes the following steps: (1) Obtain the extent of the oil slick at sea, preliminarily determine the monitoring points for oil slick monitoring, and use multiple UAVs to monitor the monitoring points; (2) With the goal of minimizing the energy consumption of multiple UAVs, and with the UAVs’ endurance, arrival time and oil film drift tolerance as constraints, a multi-UAV task allocation model is constructed. (3) Based on the location of monitoring points, UAV parameters and multi-UAV task allocation model of oil film monitoring, the task sequence of UAVs visiting monitoring points is used as the optimization variable. Combined with the objective function, the task allocation is carried out by the hybrid learning pigeon flock optimization algorithm, and the task allocation scheme is output. The hybrid learning pigeon flock optimization algorithm introduces a Q-learning mechanism in the second stage of the pigeon flock algorithm. Based on the population diversity and temporal satisfaction rate, it determines the elimination ratio and the type of eliminated individuals. The landmark operator is updated based on the surviving individuals after elimination, and the pigeon flock algorithm is performed based on the updated landmark operator.
2. The multi-UAV task allocation method for marine oil slick monitoring according to claim 1, characterized in that, The method of using multiple drones to monitor a monitoring point includes: using one fixed-wing drone and one multi-rotor drone to monitor a monitoring point, wherein the arrival time of the fixed-wing drone at the monitoring point is earlier than the arrival time of the multi-rotor drone, and the arrival time difference between the fixed-wing drone and the multi-rotor drone at the same monitoring point is within a preset threshold.
3. The multi-UAV task allocation method for marine oil slick monitoring according to claim 2, characterized in that, The multi-UAV task allocation model is as follows: ; in, Let be the objective function. Total flight distance Penalties for timing violations of arrival times for fixed-wing and multi-rotor drones. For load balancing variance, , , These are the weighting coefficients. Indicates the first monitoring points Assigned to the fixed-wing drones Assignment variables, For the number of fixed-wing drones, Indicates the first monitoring points Assigned to the Multi-rotor drones Assignment variables, For the number of monitoring points, For the number of multi-rotor drones, Represents the set of monitoring points. This indicates the locations of all drone take-off and landing bases. Indicates monitoring point The initial geographical coordinates, Indicates fixed-wing unmanned aerial vehicle The first monitoring point visited, Represents Euclidean distance. Indicates monitoring point The initial geographical coordinates, Indicates fixed-wing unmanned aerial vehicle The first visit One monitoring point, Indicates monitoring point The initial geographical coordinates, Indicates fixed-wing unmanned aerial vehicle The first visit One monitoring point, Indicates fixed-wing unmanned aerial vehicle The total number of monitoring points allocated, Indicates fixed-wing unmanned aerial vehicle Maximum range Indicates monitoring point The initial geographical coordinates, Indicates multi-rotor drone The first monitoring point visited, Indicates monitoring point The initial geographical coordinates, Indicates multi-rotor drone The first visit One monitoring point, Indicates monitoring point The initial geographical coordinates, Indicates multi-rotor drone The first visit One monitoring point, Indicates multi-rotor drone The total number of monitoring points allocated, Indicates multi-rotor drone Maximum range This indicates that you have been assigned to a monitoring point. Multi-rotor drones arrive at monitoring points At that moment, This indicates that you have been assigned to a monitoring point. The fixed-wing drone arrived at the monitoring point At that moment, The maximum allowable time difference, Indicates oil film drift tolerance. Indicates the oil film drift velocity. Indicates monitoring point exist Geographic coordinates at any given time Indicates monitoring point exist Geographic coordinates at any given time Indicates monitoring point exist Geographic coordinates at any given time This indicates the maximum number of monitoring points that each fixed-wing drone can visit in a single launch. This indicates the maximum number of monitoring points that each multi-rotor drone can visit in a single launch. Indicates fixed-wing unmanned aerial vehicle Arrival at the monitoring point At that moment, Indicates multi-rotor drone Arrival at the monitoring point At that moment.
4. The multi-UAV task allocation method for marine oil slick monitoring according to claim 3, characterized in that, The total flight distance The calculation formula is: ; in, Indicates fixed-wing unmanned aerial vehicle The geographical coordinates of the first monitoring point visited. Indicates fixed-wing unmanned aerial vehicle The first visit The geographical coordinates of each monitoring point Indicates fixed-wing unmanned aerial vehicle The first visit The geographical coordinates of each monitoring point Indicates multi-rotor drone The geographical coordinates of the first monitoring point visited. Indicates multi-rotor drone The first visit The geographical coordinates of each monitoring point Indicates multi-rotor drone The first visit The geographical coordinates of each monitoring point; The timing violation penalty for the arrival time of the fixed-wing UAV and multi-rotor UAV The calculation formula is: ; The load balancing variance The calculation formula is: ; in, For fixed-wing unmanned aerial vehicles Total flight distance This represents the average flight distance of all fixed-wing UAVs. For multi-rotor drones Total flight distance This represents the average flight distance of all multi-rotor drones.
5. The multi-UAV task allocation method for marine oil slick monitoring according to claim 3, characterized in that, The steps for task allocation using the hybrid learning pigeon flock optimization algorithm include: Obtain the initial position and initial speed of the pigeons, and determine the initial population based on the initial position and initial speed; The fitness value of each pigeon in the initial population is calculated using the objective function, and the position of the target pigeon is determined based on the fitness value of each pigeon. The population diversity and temporal constraint satisfaction rate are also calculated. Based on the location of the target pigeon, the position and speed of each pigeon are updated using a geomagnetic operator; After the number of iterations reaches the first value, the current state is determined based on population diversity and temporal constraint satisfaction rate. An action is selected based on the Q value. The action includes elimination ratio and elimination individual type. The elimination ratio determines the individual with the worst fitness in the current population to be eliminated. The elimination individual type determines whether to eliminate directly or replace the eliminated individual. Direct elimination means directly deleting the eliminated individual and randomly copying it from the surviving individuals to supplement the original population size. Replacement elimination means randomly selecting two parents from the surviving individuals, performing single-point crossover between the two parents to generate two offspring, and randomly selecting one of the two offspring to replace the individual to be eliminated. The action includes: directly eliminating or replacing the determined individual to be eliminated. The landmark operator is updated based on the center position of the surviving individuals after the action is performed; the position of each pigeon is updated based on the updated landmark operator; the fitness value of each pigeon is calculated; and the global optimal fitness is updated. After performing an action, observe the new state and calculate the reward, then update the Q value; If the iteration stopping condition is met, output the position of the pigeon with the global optimal fitness, and determine the task allocation scheme based on the pigeon position.
6. The multi-UAV task allocation method for marine oil slick monitoring according to claim 5, characterized in that, The updated landmark operator updates the position of each pigeon as follows: ; in, for The position vector of the pigeon at any given moment. for The position vector of the pigeon at any given moment. The numbers are uniformly random. The center position of the surviving individual after the action was performed: ; in, Indicates surviving individuals, The maximum fitness value among surviving individuals. For the first The position vectors of the pigeons. For the first The fitness value of each pigeon.
7. The multi-UAV task allocation method for marine oil slick monitoring according to claim 5, characterized in that, The auction mechanism is used to convert pigeon locations into task allocation schemes. The steps for determining the task allocation scheme based on pigeon locations are as follows: Position the pigeon in the middle front Each component corresponds to a bid price for a fixed-wing UAV; later Each component corresponds to a bid price for a multi-rotor drone; All fixed-wing UAVs are sorted from highest to lowest bid price to obtain the first ordered list; all multi-rotor UAVs are sorted from highest to lowest bid price to obtain the second ordered list; and each monitoring point is sorted from highest to lowest oil film thickness or urgency to obtain the third ordered list. Based on the sorting of the third ordered list, combined with the multi-UAV task allocation model, fixed-wing UAVs and multi-rotor UAVs are selected from the first ordered list and the second ordered list for each monitoring point in turn; and the remaining endurance and number of assigned tasks of each fixed-wing UAV and multi-rotor UAV are updated. For each drone, its assigned monitoring points are arranged from closest to furthest from the base to obtain the access order, which is the task allocation scheme.
8. The multi-UAV task allocation method for marine oil slick monitoring according to claim 7, characterized in that, The task allocation scheme is partially modified to ensure that the arrival time difference between fixed-wing UAVs and multi-rotor UAVs at each monitoring point meets a preset constraint; the partial modification actions include: Action 1: Insert a virtual detour point between the target monitoring point and the previous monitoring point in the drone's access sequence. The virtual detour point is located on the perpendicular bisector of the line connecting the previous monitoring point and the target monitoring point, thereby increasing the drone's flight distance, and ensuring that the total flight distance after detour does not exceed its remaining endurance. Action 2: Move the target monitoring point forward in the drone's access sequence, or remove unnecessary detour points in the drone's path; Action 3: Transfer the monitoring task of the target monitoring point to other drones that are idle or have a task load less than the threshold; The partial repair process is as follows: If the arrival order of the fixed-wing drone and the multi-rotor drone is reversed, the fixed-wing drone will perform action 2. If this is ineffective, the multi-rotor drone will perform action 1. If this is ineffective, either the fixed-wing drone or the multi-rotor drone will perform action 3. If the arrival time difference between the fixed-wing drone and the multi-rotor drone is greater than the maximum allowable time difference, the multi-rotor drone will perform action 2. If this is ineffective, the fixed-wing drone will perform action 1. If this is also ineffective, either the fixed-wing drone or the multi-rotor drone will perform action 3. If the arrival time difference between the fixed-wing UAV and the multi-rotor UAV meets the preset constraints after the repair, the repaired access order will be used as the new task allocation scheme; otherwise, the original task allocation scheme will be retained.
9. The multi-UAV task allocation method for marine oil slick monitoring according to claim 8, characterized in that, It also includes step (4), every interval After a few seconds, the geographical coordinates of each monitoring point that the drone has not reached are determined again. If the difference between the geographical coordinates of the monitoring point at this time and the geographical coordinates at the previous time is greater than the tolerance threshold, the monitoring point is marked as invalid and the invalid monitoring point in the drone access sequence is released. Step (3) is repeated for the failed monitoring points and idle drones to obtain the redistribution results. The redistribution results are then merged with the original task allocation scheme to form a new task allocation scheme. After each redistribution, the oil film drift speed is updated.
10. A multi-UAV task allocation system for marine oil slick monitoring, characterized in that, include: The data acquisition module is used to obtain the extent of the oil slick at sea, preliminarily determine the monitoring points for oil slick monitoring, and use multiple UAVs to monitor the monitoring points; The model building module is used to construct a multi-UAV task allocation model with the objective function of minimizing the energy consumption of multiple UAVs and the constraints of UAV endurance, arrival time and oil film drift tolerance. The task allocation module is used to allocate tasks based on the location of monitoring points, UAV parameters and multi-UAV task allocation model of oil film monitoring. It takes the task sequence of UAVs visiting monitoring points as optimization variables, combines the objective function and uses the hybrid learning pigeon flock optimization algorithm to allocate tasks and output the task allocation scheme. The hybrid learning pigeon flock optimization algorithm introduces a Q-learning mechanism in the second stage of the pigeon flock algorithm. Based on the population diversity and temporal satisfaction rate, it determines the elimination ratio and the type of eliminated individuals. The landmark operator is updated based on the surviving individuals after elimination, and the pigeon flock algorithm is performed based on the updated landmark operator.