Multi-UAV Three-Dimensional Path Planning Method for Task Sets

By constructing an integer planning model and combining the Q-learning mechanism genetic algorithm, the problems of low patrol efficiency and high labor cost in drone inspections are solved, and the three-dimensional path planning of multi-drone is realized, which improves the intelligence level of drone inspections.

CN116360481BActive Publication Date: 2025-07-22HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310136735.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-10
Publication Date
2025-07-22
Estimated Expiration
2043-02-10

AI Technical Summary

Technical Problem

Traditional drone inspection methods have problems such as low patrol efficiency, high labor costs, and solidified inspection routes.

Method used

A integer planning model is constructed to minimize the time when all drones complete patrol tasks as the optimization goal, and to solve it in combination with the genetic algorithm of the Q-learning mechanism to plan the three-dimensional path of multiple drones.

Benefits of technology

It has achieved reasonable selection of task points based on actual inspection task needs, improved drone inspection efficiency, reduced labor costs, and promoted the in-depth application of drones in the fields of power inspection, hydropower station inspection, urban inspection and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116360481B_ABST
    Figure CN116360481B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-UAV three-dimensional path planning method, system, storage medium and electronic device for a task set, which relates to the technical field of UAV inspection. In the present invention, with the goal of minimizing the time for all UAVs to complete the inspection tasks, an integer programming model is constructed, and a genetic algorithm integrating the Q-learning mechanism is used to solve the model, so as to realize how to reasonably select task points from the task point set according to the actual inspection task requirements and then plan the multi-UAV three-dimensional path. It has important theoretical value and practical significance for promoting the in-depth application of UAV inspection in fields such as power inspection, hydropower station inspection, and urban inspection, and improving the intelligent inspection level of UAVs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of UAV inspection, and particularly relates to a multi-UAV three-dimensional path planning method, system, storage medium and electronic device for a task set. Background Art

[0002] In recent years, with the rapid development of UAVs, UAVs have become one of the important operation methods for data collection in fields such as urban inspection, transmission line inspection, and hydropower station inspection.

[0003] Traditional inspection operations mainly rely on manual work, and there are certain blind spots in terms of inspection efficiency, accuracy, and safety based on experience. At present, the inspection mode is gradually changing from manual inspection to human-machine collaboration. The UAV inspection operation method has become a mainstream method in various inspection fields and has been widely applied to daily inspection work. At the same time, the UAV inspection method has evolved from manual control of UAV flight by personnel to the stage of automated inspection. By manually setting the flight route of the UAV in advance, the inspection process is further simplified and the inspection efficiency is improved.

[0004] However, the existing inspection methods have problems such as low inspection efficiency, high labor cost, and fixed inspection routes. Summary of the Invention

[0005] (I) Technical Problems to be Solved

[0006] Aiming at the deficiencies of the prior art, the present invention provides a multi-UAV three-dimensional path planning method, system, storage medium and electronic device for a task set, which solves the technical problems of low inspection efficiency, high labor cost, and fixed inspection routes existing in traditional inspection methods.

[0007] (II) Technical Solutions

[0008] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0009] A multi-UAV three-dimensional path planning method for a task set includes:

[0010] S1. Obtain the task set to be inspected and UAV resources;

[0011] S2. According to the task set to be inspected and UAV resources, construct an integer programming model with the goal of minimizing the time for all UAVs to complete the inspection tasks;

[0012] S3. Solve the integer programming model by using a genetic algorithm integrating the Q-learning mechanism to obtain the multi-UAV three-dimensional path planning result.

[0013] Preferably, the integer programming model in S2 includes:

[0014] The objective function with the goal of minimizing the time for all drones to complete the inspection task:

[0015]

[0016] where i and j are indices of task points or virtual sites, i, j ∈ N ∪ D;

[0017] N is the set of task points;

[0018] D is the set of virtual sites. One drone corresponds to one virtual site. The virtual sites are used to ensure that each drone starts from the starting point and finally returns to the starting point. The coordinates of the virtual sites are all the coordinates of the starting point;

[0019] R ij is the reliable distance between task point i and task point j;

[0020] F l is the set of task points corresponding to cluster l;

[0021] m l is the number of task points to be visited in cluster l;

[0022] v is the speed of the drone;

[0023] x ij is a decision variable. When task point i is connected to task point j, it takes 1, otherwise it takes 0.

[0024] Preferably, the integer programming model in S2 further includes:

[0025] Constraint conditions:

[0026] (1) Each virtual site enters and exits once, that is, each drone starts from the starting point and then returns to the starting point,

[0027]

[0028]

[0029] (2) If there is an incoming flow at point j, then point j is visited by the drone,

[0030]

[0031] (3) If there is an incoming flow at point i, then point i is also visited by the drone,

[0032]

[0033] (4) A specified number of task points in each cluster need to be visited,

[0034]

[0035] (5) Each point can be visited at most once.

[0036]

[0037] (6) If the path (i, j) is selected, then points i and j should be visited by the same drone.

[0038]

[0039] (7) Eliminate the sub - circuit constraint.

[0040]

[0041] where k is the drone index, k ∈ K;

[0042] n is the sum of the number of task points and virtual sites;

[0043] z ik and z jk respectively represent the order in which task points i and j are visited by drone k;

[0044] y ik and y jk are both decision variables. When drone k visits task point i, y ik takes 1, otherwise 0; when drone k visits task point j, y jk takes 1, otherwise 0.

[0045] Preferably, the S3 includes:

[0046] S31. Let the iteration number t = 1; generate an initial population according to the constraint conditions of the model and initialize the Q - table. Among them, the chromosome encoding rule specifically means:

[0047] A single three - line chromosome represents a solution. The total length of the chromosome is equal to the number of task points. Each task point is pre - marked with a unique serial number. The second - line chromosome represents the task - point sequence. The first - line chromosome represents the cluster sequence corresponding to the task - point sequence. The third - line chromosome represents the drone number, indicating that the task point is selected by this drone for visit, and if it is 0, it is not visited;

[0048] S32. If the maximum iteration number is reached, stop and decode the current optimal chromosome as the multi - drone three - dimensional path planning result, otherwise continue to execute;

[0049] S33. Calculate the objective function value of the population. Among them, the linear transformation method is used to transform the objective function of the model into a fitness function, and the transformation formula is as follows:

[0050] F′ = aF + b

[0051] Among them, F′ represents the fitness function, a and b are hyperparameters of the linear equation, and a < 0;

[0052] S34. Evaluate the state of the current population according to the population objective function value; select actions according to the state of the current population, calculate the reward, and update the Q table;

[0053] S35. Select M parents from the current population using the roulette wheel strategy;

[0054] S36. Perform crossover operations on the M parents according to the crossover probability determined by the action selected in S34;

[0055] S37. Perform mutation operations on the M individuals after crossover according to the mutation probability determined by the action selected in S34;

[0056] S38. Combine the parent population and the offspring population to obtain a population of size 2M;

[0057] S39. Select M individuals from the 2M population using the reinsertion method based on fitness ranking to obtain a new generation population; let t = t + 1, and return to S32.

[0058] Preferably, in S34, evaluating the state of the current population according to the population objective function value specifically includes:

[0059] S341. Solve the average population objective function z1, the diversity z2 of the population, and the best objective function value z3 of the population of the current population respectively;

[0060]

[0061]

[0062]

[0063] Among them, is the objective function value of the u-th individual in the t-th iteration; is the objective function value of the w-th individual in the 1st iteration; is the objective function value of the optimal individual in the t-th iteration, is the objective function value of the worst individual in the t-th iteration;

[0064] S342. Solve the state of the current population;

[0065] S * = w1*z1 + w2*z2 + w3*z3

[0066] Among them, w1, w2, and w3 are the weights of the corresponding measurement values respectively.

[0067] Preferably, in S34, actions are selected according to the current state of the population, rewards are calculated, and the Q-table is updated, specifically including:

[0068] S343. According to the current state of the population, the ε-greedy strategy is used to select actions, and the crossover probability and mutation probability are determined; the ε-greedy strategy means:

[0069]

[0070] where π(s t , a t ) represents the optimal strategy selected in state s t ;

[0071] Define S = {s1,..., sn} as the state set, where n is the total number of states; A = {a1,..., am} as the set of candidate actions to be executed, where m is the total number of actions, and a1 = (Pc1, Pm1), a2 = (Pc2, Pm2), and so on; Pc is the crossover probability, and Pm is the mutation probability; Q(s t , a) represents the set of actions that can be selected in state s t ; Max a (Q(s t , a)) represents the action with the largest Q value in the Q-table selected in state s t ; ξ is a random number; rand(0, 1) represents a decimal randomly generated between 0 and 1; ε represents the exploration rate parameter;

[0072] S344. Calculate the reward,

[0073] r = y1 * r1 + y2 * r2

[0074]

[0075]

[0076] where r is the reward function value, y1 and y2 are the weights of the corresponding metric values respectively; r1 > 0 means that if the objective function value of the optimal individual in the t-th generation is better than that in the (t - 1)-th generation, the current crossover probability Pc is rewarded; r2 > 0 means that if the average objective function value in the t-th generation is better than that in the (t - 1)-th generation, the current mutation probability Pm is rewarded;

[0077] S345. Update the Q-table,

[0078] Q(s t , a t ) = (1 - α)Q(s t , a t)+α(r+γmax a Q(s t+1 ,a t+1 ))

[0079]

[0080] Among them, Q(s t ,a t ) means in state s t Take action t The obtained Q value, Q(s t+1 ,a t+1 ) means in state s t+1 Take action t+1 The obtained Q value; α is the learning factor of Q-learning; r is the current state s t The reward under ; γ is the discount factor of the algorithm.

[0081] A multi-UAV three-dimensional path planning system for a task set, comprising:

[0082] Data acquisition module, used to obtain the task set and drone resources to be inspected;

[0083] A model building module is used to build an integer programming model based on the set of tasks to be inspected and the drone resources, with the optimization goal of minimizing the time for all drones to complete the inspection tasks;

[0084] The result solving module is used to solve the integer programming model by using a genetic algorithm integrating a Q-learning mechanism to obtain the three-dimensional path planning results of multiple UAVs.

[0085] A storage medium stores a computer program for three-dimensional path planning of multiple unmanned aerial vehicles for a task set, wherein the computer program enables a computer to execute the three-dimensional path planning method of multiple unmanned aerial vehicles for a task set as described above.

[0086] An electronic device, comprising:

[0087] one or more processors;

[0088] Memory; and

[0089] One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include a method for executing a multi-UAV three-dimensional path planning method for a task set as described above.

[0090] (III) Beneficial effects

[0091] The present invention provides a three-dimensional path planning method, system, storage medium and electronic device for a multi-UAV facing a task set. Compared with the prior art, the following beneficial effects are achieved:

[0092] In the present invention, taking the shortest time for all UAVs to complete the inspection tasks as the optimization goal, an integer programming model is constructed, and a genetic algorithm integrating the Q-learning mechanism is used to solve the model, so as to realize how to reasonably select task points from the task point set according to the actual inspection task requirements and then plan the three-dimensional paths of multiple UAVs. It has important theoretical value and practical significance for promoting the in-depth application of UAV inspection in the fields of power inspection, hydropower station inspection, urban inspection, etc., and improving the intelligent inspection level of UAVs. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0094] Figure 1 An example provided by an embodiment of the present invention with 4 clusters, 11 task points and 2 UAVs (the cube is a building);

[0095] Figure 2 A block diagram of a three-dimensional path planning method for a multi-UAV facing a task set provided by an embodiment of the present invention;

[0096] Figure 3 A flowchart of a genetic algorithm integrating the Q-learning mechanism provided by an embodiment of the present invention;

[0097] Figure 4 An example of a chromosome coding rule provided by an embodiment of the present invention;

[0098] Figure 5 A schematic diagram of the roulette wheel selection method provided by an embodiment of the present invention;

[0099] Figure 6 A schematic diagram of the chromosome crossover process provided by an embodiment of the present invention;

[0100] Figure 7 A schematic diagram of the chromosome mutation process provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0101] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0102] By providing a multi-UAV three-dimensional path planning method, system, storage medium, and electronic device for a task set, the embodiments of the present application solve the technical problems of low inspection efficiency, high labor cost, and fixed inspection routes in traditional inspection methods, realize intelligent collaborative inspection of multiple UAVs, replace the traditional inspection method, and complete the reform of the UAV inspection mode.

[0103] The general idea of the technical solutions in the embodiments of the present application to solve the above technical problems is as follows:

[0104] As proposed in the background technology, studying how to deploy multiple UAVs to collaboratively complete the rapid inspection of task points can greatly reduce manual repetitive work and reduce the impact of the number of staff on the inspection work. To solve these problems, the mode of multi-UAV collaborative inspection is becoming a new direction for power inspection. In addition, a common problem in various inspection scenarios is that usually, for each key point to be inspected, it can be photographed from multiple different positions, that is, each key point to be inspected corresponds to multiple inspection task points. When planning the UAV path, studying how to select task points at appropriate positions from the task point set according to the actual inspection task requirements and the number of task points that need to be visited for each key point can further reduce the inspection time of the UAV, improve the inspection efficiency, realize intelligent collaborative inspection of multiple UAVs, replace the traditional inspection method, and complete the reform of the UAV inspection mode.

[0105] To address the above problems, the embodiments of the present invention intend to study the multi-UAV three-dimensional path planning technology for a task point set. With the goal of minimizing the time for all UAVs to complete the inspection task, an integer programming model is constructed, and a genetic algorithm integrated with the Q-learning mechanism is used to solve the model, so as to realize how to reasonably select task points from the task point set according to the actual inspection task requirements and then plan the multi-UAV three-dimensional path. This technology has important theoretical value and practical significance for promoting the in-depth application of UAV inspection in fields such as power inspection, hydropower station inspection, and urban inspection, and improving the intelligent inspection level of UAVs.

[0106] Before formally introducing the technical solutions provided by the embodiments of the present invention, it is necessary to provide a description of the problems:

[0107] There are n buildings and l key points to be inspected in the same area, and it is required to use k drones to complete the inspection task. Each key point corresponds to a cluster, and the cluster contains F l shooting task points with different positions. For the l-th cluster, the drone visits m l (m l ≤F l ) task points to complete the inspection of the key point. All drones start from the starting point, return to the starting point after completing the inspection task, and a reasonable drone inspection route needs to be obtained to minimize the longest inspection time of all drones.

[0108] Figure 1 Fig. shows an example with 2 buildings, 4 clusters, and 11 task points. Clusters 1, 2, and 3 each have three task points, and cluster 4 has 2 task points. In this example, one task point is specified to be visited in each cluster, and two drones need to start from the starting point and finally return to the starting point. To clarify the scope of application of this article, the following assumptions are made: The drones fly at an average speed, and the time consumed when the drones reach the task points to photograph the components is ignored.

[0109] To better understand the above technical solution, the above technical solution will be described in detail below in combination with the accompanying drawings of the specification and specific implementation manners.

[0110] Embodiment:

[0111] As Figure 2 shown, an embodiment of the present invention provides a multi-drone three-dimensional path planning method for a task set, including:

[0112] S1. Obtain the task set to be inspected and drone resources;

[0113] S2. According to the task set to be inspected and drone resources, construct an integer programming model with the goal of minimizing the time for all drones to complete the inspection task;

[0114] S3. Solve the integer programming model by using a genetic algorithm integrating the Q-learning mechanism to obtain the multi-drone three-dimensional path planning result.

[0115] The embodiment of the present invention has important theoretical value and practical significance for promoting the in-depth application of drone inspection in fields such as power inspection, hydropower station inspection, and urban inspection, and improving the intelligent inspection level of drones.

[0116] Next, each step of the technical solution will be introduced:

[0117] In step S1, obtain the task set to be inspected and drone resources.

[0118] To facilitate the description of the model, the symbols of the mathematical model of MinMax MDFTSP (multi - depot family traveling salesman problem) related to the embodiments of the present invention are defined as follows:

[0119] Decision variables:

[0120]

[0121]

[0122] Symbol definitions:

[0123]

[0124] Parameter definitions:

[0125] <![CDATA[R ij > Reliable distance from task point i to task point j <![CDATA[F l > Set of task points corresponding to cluster l <![CDATA[m l > Number of task points to be visited in cluster l v UAV speed n Sum of the number of task points and virtual sites <![CDATA[z ik > Order in which task point i is visited by UAV k

[0126] In step S2, an integer programming model is constructed with the objective of minimizing the time for all drones to complete the inspection tasks, in combination with the reliable distance.

[0127] The integer programming model includes:

[0128] The objective function with the objective of minimizing the time for all drones to complete the inspection tasks:

[0129]

[0130] Constraint conditions:

[0131] (1) Each virtual site is entered and exited once, that is, each drone starts from the starting point and then returns to the starting point.

[0132]

[0133]

[0134] (2) If there is an inflow at point j, then point j is visited by the drone.

[0135]

[0136] (3) If there is an inflow at point i, then point i is also visited by the drone.

[0137]

[0138] (4) A specified number of task points in each cluster need to be visited.

[0139]

[0140] (5) Each point can be visited at most once,

[0141]

[0142] (6) If the path (i, j) is selected, then points i and j should be visited by the same UAV,

[0143]

[0144] (7) Eliminate the sub - circuit constraint,

[0145]

[0146] In step S3, a genetic algorithm integrating the Q - learning mechanism is used to solve the integer programming model, and the multi - UAV three - dimensional path planning result is obtained.

[0147] As a global search method, the genetic algorithm (GA) can find an approximate optimal solution to the problem in a relatively short time by mimicking the "survival of the fittest" in the biological world and can well solve the TSP problem. The hyperparameters in the genetic algorithm have an important impact on its optimization performance. How to intelligently determine its hyperparameters during the optimization process is a confusing problem. This always requires a large number of repeated experiments to determine, and once the parameters are set, their values will not change throughout the optimization process. Adjusting the hyperparameters through preset rules is a method that can meet the above requirements, but the performance of this method is not satisfactory and the generalization ability is poor.

[0148] An embodiment of the present invention proposes a new framework. During the iteration process of the genetic algorithm, the Q - learning mechanism is combined to intelligently change and adjust the hyperparameters of the genetic algorithm, that is, the probability of chromosome crossover and mutation, and QLGA is obtained to solve the problem. The difference from the conventional genetic algorithm is that in the algorithm, after each new population is generated, by evaluating the quality of the current population, the crossover and mutation probability of the next - step chromosome is dynamically adjusted according to the state of the current population, so that the algorithm can make a better balance between exploration and exploitation at each stage.

[0149] As Figure 3 shown, the overall algorithm flow is as follows:

[0150] S31. Let the iteration number t = 1; generate an initial population according to the constraint conditions of the model and initialize the Q - table; where, as Figure 4 shown, the chromosome encoding rule specifically refers to:

[0151] A single three - line chromosome represents a solution. The total length of the chromosome is equal to the number of task points. Each task point is pre - marked with a unique serial number. The second - line chromosome represents the task - point sequence. The first - line chromosome represents the cluster sequence corresponding to the task - point sequence, that is, the set where the task point is located. The third - line chromosome represents the UAV serial number, indicating that the task point is selected by this UAV for visit. If it is 0, it is not visited.

[0152] Based on the aforementioned model, during the process of initializing the chromosome population: First, generate a sequence of numbers from small to large according to the number of task points. The cluster serial number corresponding to each task point, which is known information, is also generated accordingly. When generating the third - line chromosome, it is necessary to ensure that the number of visited task points in each cluster reaches the set quantity requirement. Finally, shuffle the current chromosome matrix column - by - column and repeat multiple times to obtain the initialized population.

[0153] In addition, in the Q - learning algorithm, the Q - table is used to store the learning experience of the agent in different states, where the number of rows and columns of the table is equal to the defined states and actions. All values of the Q - table at the initial moment are equal to zero. This means that the algorithm has no experience to use, and action selection can be randomly performed through the Q - learning algorithm. See the subsequent steps for specific content.

[0154] S32. If the maximum number of iterations is reached, stop and decode the current optimal chromosome as the multi - UAV three - dimensional path planning result; otherwise, continue to execute.

[0155] S33. Calculate the population objective - function value. Among them, the linear transformation method is used to transform the objective function of the model into a fitness function. At this time, the smaller the value of the objective - function value, the larger the fitness value, and the better the represented solution. The transformation formula is as follows:

[0156] F′ = aF + b

[0157] where F′ represents the fitness function, a and b are hyperparameters of the linear equation, and a < 0;

[0158] And F′ should satisfy the following two conditions:

[0159] (1) To ensure that the expected replication number of individuals with fitness at the average value in the next generation is 1, it is necessary to make the average value of the original fitness Favg = the average value of the new fitness F′avg;

[0160] (2) To control the replication number of individuals with the maximum fitness in the next generation, it is set that the maximum value of the transformed fitness is equal to a specified multiple of the average value of the original fitness. For example: F′max = 2 * Favg.

[0161] S34. Evaluate the current population state based on the population objective function value; select actions according to the current population state, calculate the reward, and update the Q-table.

[0162] Reinforcement learning is a subset of artificial intelligence algorithms. The reinforcement learning framework consists of five main components, namely ① the agent, ② the environment, ③ the state, ④ the action, and ⑤ the reward. In this process, the agent interacts with the environment and is trained by evaluating actions according to environmental signals. By seeking to select appropriate actions, the best strategy is found to maximize the long-term reward.

[0163] Embed the Q-learning algorithm into the genetic algorithm. Determine the population state based on the objective function value of the current population. By setting the reward, continuously update the value of the Q-table to dynamically adjust the two main parameters of the genetic algorithm: the crossover rate Pc and the mutation probability Pm. These parameters have a significant impact on the performance of the GA and help the genetic algorithm make a better trade-off between exploration and exploitation at each stage.

[0164] Specifically, in this step, evaluating the current population state based on the population objective function value specifically includes:

[0165] S341. Solve the average population objective function z1, the population diversity z2, and the best population objective function value z3 of the current population respectively;

[0166]

[0167]

[0168]

[0169] where is the objective function value of the u-th individual in the t-th iteration; is the objective function value of the w-th individual in the 1st iteration; is the objective function value of the optimal individual in the t-th iteration, is the objective function value of the worst individual in the t-th iteration;

[0170] S342. Solve the current population state;

[0171] S * = w1*z1 + w2*z2 + w3*z3

[0172] where w1, w2, and w3 are the weights of the corresponding metric values respectively. Specifically, the possible values of S* can be divided into 10 states according to the interval: [0, 0.1], [0.1, 0.2]…[0.9, 1.0].

[0173] And in this step, action selection is performed according to the state of the current population, the reward is calculated, and the Q-table is updated, specifically including:

[0174] S343. According to the state of the current population, an ε-greedy strategy is adopted for action selection to determine the crossover probability and mutation probability.

[0175] The ε-greedy strategy can balance between exploration and exploitation in the Q-learning algorithm, specifically referring to:

[0176]

[0177] where π(s t ,a t ) represents the optimal strategy selected in state s t ;

[0178] Define S = {s1,..., sn} as the set of states, where n is the total number of states;

[0179] A = {a1,..., am} is the set of candidate actions to be executed, where m is the total number of actions. Among them, a1 = (Pc1, Pm1), a2 = (Pc2, Pm2), and so on. Each time an iteration is performed, a pair of parameters is selected from the action set to adjust the crossover rate and mutation rate of the population in the genetic algorithm; that is, Pc is the crossover probability and Pm is the mutation probability. In particular, the crossover P c ∈[0.4, 0.9), with an interval value of 0.05; the mutation P m ∈[0.01, 0.21), with an interval value of 0.02;

[0180] Q(s t ,a) represents the set of actions that can be selected in state s t ; Max a (Q(s t ,a)) represents the action with the largest Q value in the Q-table selected in state s t ; ξ is a random number; rand(0,1) represents generating a random decimal between 0 and 1; ε represents the exploration rate parameter.

[0181] S344. Calculate the reward,

[0182] r = y1 * t1 + y2 * r2

[0183]

[0184]

[0185] Among them, r is the reward function value, and y1 and y2 are the weights of the corresponding metric values respectively; r1>0 indicates that if the objective function value of the optimal individual in the t-th generation is better than that in the (t-1)-th generation, the current crossover probability Pc is rewarded as effective; r2>0 indicates that if the average objective function value in the t-th generation is better than that in the (t-1)-th generation, the current mutation probability Pm is rewarded as effective.

[0186] S345. Update the Q-table;

[0187] Q(s t ,a t ) = (1 - α)Q(s t ,a t ) + α(r + γmax a Q(s t+1 ,a t+1 ))

[0188]

[0189] Among them, Q(s t ,a t ) represents the Q-value obtained by taking action a t under state s t , and Q(s t+1 ,a t+1 ) represents the Q-value obtained by taking action a t+1 under state s t+1 ; α is the learning factor of Q-learning, which is used to determine the severity of the algorithm's response to rewards. The larger the α value, the greater the fluctuation of the Q-table in each iteration; r is the reward under the current state s t ; γ is the discount coefficient of the algorithm, which balances the effects of current and future rewards. The higher the γ value, the more future rewards are prioritized.

[0190] S35. Select M parents from the current population using the roulette wheel strategy;

[0191] As Figure 5 shown, the individuals in the population are mapped to continuous segments of the interval, and the length of the segment where each individual is located is proportional to its fitness. Generate a random number, select the corresponding individual according to the segment it falls into, and repeat this process until the required number of individuals is obtained.

[0192] S36. Perform crossover operations on the M parents according to the crossover probability determined by the action selected in S34;

[0193] The crossover operation changes the gene sequence of chromosomes by crossing gene segments between chromosomes, which can increase the population diversity and improve the global search ability of the genetic algorithm. Here, the first row of the chromosome represents the cluster information of the task points and does not participate in crossover and mutation, but only changes along with the second row of the chromosome. The second row of the chromosome performs the crossover operation using the partially-matched crossover method, as Figure 6 shown.

[0194] By randomly selecting two crossover points to determine the crossover region, generally two invalid chromosomes will be obtained after crossover, and individual genes may appear repeatedly. To repair the chromosomes, establish the matching relationship of each chromosome within the crossover region, and then apply this matching relationship to the duplicate genes outside the crossover region to eliminate conflicts. The third row of the chromosome uses the two-point crossover method. First, randomly determine the gene positions to be crossed, and exchange the selected segments of these two groups of genes to obtain two offspring chromosomes.

[0195] S37. Perform the mutation operation on the M individuals after crossover according to the mutation probability determined by the action selected in S34;

[0196] The mutation operation generates new chromosomes by changing genes or gene positions in the chromosomes, increasing the population diversity to avoid the algorithm falling into local optima. As Figure 7 shown, here the second row of the chromosome uses segment inversion mutation. By randomly selecting two mutation points to determine the mutation region, and then reversing the chromosome segment in the mutation region. The third row of the chromosome uses the mutation method of integer value mutation. First, judge whether the task point allocation scheme in the third row meets the requirement of the number of task points that must be visited in each cluster. If not, mutate the numbers in that cluster.

[0197] S38. Combine the parent population and the offspring population to obtain a population of size 2M;

[0198] S39. Select M individuals from the 2M population using the reinsertion method based on fitness ranking to obtain a new generation population; let t = t + 1, and return to S32;

[0199] First, calculate the fitness of all individuals in the parent population and the new population obtained by crossover and mutation respectively. For example, use the first 90% of the individuals in the new population to replace the last 90% of the individuals in the parent population to update the population and obtain the offspring population. This can ensure that the optimal individuals are inherited to the next generation and enable the algorithm to have better global search ability and avoid falling into local optima.

[0200] The embodiment of the present invention provides a multi-UAV three-dimensional path planning system for a task set, including:

[0201] A data acquisition module for acquiring the task set to be inspected and UAV resources;

[0202] A model construction module, configured to construct an integer programming model with the goal of minimizing the time for all drones to complete inspection tasks, according to the task set to be inspected and drone resources.

[0203] A result solving module, configured to solve the integer programming model by using a genetic algorithm integrating the Q-learning mechanism, and obtain a multi-drone three-dimensional path planning result.

[0204] An embodiment of the present invention provides a storage medium, characterized in that it stores a computer program for multi-drone three-dimensional path planning for a task set, wherein the computer program enables a computer to execute the multi-drone three-dimensional path planning method for a task set as described above.

[0205] An embodiment of the present invention provides an electronic device, including:

[0206] One or more processors;

[0207] A memory; and

[0208] One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include those for executing the multi-drone three-dimensional path planning method for a task set as described above.

[0209] It can be understood that the multi-drone three-dimensional path planning system, storage medium, and electronic device for a task set provided by the embodiments of the present invention correspond to the multi-drone three-dimensional path planning method for a task set provided by the embodiments of the present invention. For the explanations, examples, beneficial effects, etc. of the relevant content, reference can be made to the corresponding parts in the multi-drone three-dimensional path planning method for a task set, which will not be elaborated here.

[0210] In summary, compared with the prior art, the following beneficial effects are achieved:

[0211] 1. In the embodiments of the present invention, with the goal of minimizing the time for all drones to complete inspection tasks, an integer programming model is constructed, and a genetic algorithm integrating the Q-learning mechanism is used to solve the model, realizing how to reasonably select task points from the task point set according to the actual inspection task requirements and then plan a multi-drone three-dimensional path. It has important theoretical value and practical significance for promoting the in-depth application of drone inspection in fields such as power inspection, hydropower station inspection, and urban inspection, and improving the intelligent inspection level of drones.

[0212] 2. The embodiment of the present invention proposes a new framework. During the iterative process of the genetic algorithm, the Q-learning mechanism is combined to intelligently change and adjust the hyperparameters of the genetic algorithm, that is, the probabilities of chromosome crossover and mutation, so as to obtain the QLGA to solve the problem. The difference from the conventional genetic algorithm lies in that in the algorithm, after each new population is generated, by evaluating the superiority and inferiority of the current population, the crossover and mutation probabilities of the next chromosome are dynamically adjusted according to the state of the current population, so that the algorithm can make a better balance between exploration and development at each stage.

[0213] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0214] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A three-dimensional path planning method for multiple unmanned aerial vehicles oriented to a task set, characterized in that, Including: S1. Obtain the task set to be inspected and the UAV resources; S2. Based on the task set to be inspected and the UAV resources, construct an integer programming model with the goal of minimizing the time for all UAVs to complete the inspection tasks; S3. Solve the integer programming model using a genetic algorithm integrated with the Q-learning mechanism to obtain the multi-UAV three-dimensional path planning result; The S3 includes: S31. Let the iteration number t = 1; generate an initial population according to the constraint conditions of the model and initialize the Q-table; where, the chromosome encoding rule specifically refers to: A single three-line chromosome represents a solution, and the total length of the chromosome is equal to the number of task points; each task point is pre-marked with a unique serial number, the second line of the chromosome represents the task point sequence; the first line of the chromosome represents the cluster sequence corresponding to the task point sequence; the third line of the chromosome represents the UAV number, indicating that the task point is selected by this UAV for access, and if it is 0, it is not accessed; S32. If the maximum iteration number is reached, stop and decode the current optimal chromosome as the multi-UAV three-dimensional path planning result, otherwise continue; S33. Calculate the objective function value of the population; where, the objective function of the model is transformed into a fitness function using the linear transformation method, and the transformation formula is as follows: F′ = aF + b where, F′ represents the fitness function, a and b are hyperparameters of the linear equation, and a < 0, F is the objective function; S34. Judge the state of the current population according to the objective function value of the population; select actions according to the state of the current population, calculate the reward, and update the Q-table; S35. Select M parents from the current population using the roulette wheel strategy; S36. Perform crossover operations on the M parents according to the crossover probability determined by the actions selected in S34; S37. Perform mutation operations on the M individuals after crossover according to the mutation probability determined by the actions selected in S34; S38. Combine the parent population and the offspring population to obtain a population of size 2M; S39. Select M individuals from the 2M population using the reinsertion method based on fitness sorting to obtain a new generation population; let t = t + 1, and return to S32; In the S34, judging the state of the current population according to the objective function value of the population specifically includes: S341. Solve the average population objective function z1, the population diversity z2, and the best objective function value z3 of the current population respectively; Among them, is the objective function value of the $u$-th individual in the $t$-th iteration; is the objective function value of the $w$-th individual in the 1st iteration; is the objective function value of the optimal individual in the $t$-th iteration, is the objective function value of the worst individual in the $t$-th iteration; S342. Solve the state of the current population; S * = w1*z1 + w2*z2 + w3*z3 where, w1, w2, and w3 are the weights of the corresponding measurement values respectively.

2. The multi-UAV three-dimensional path planning method for a task set according to claim 1, wherein, The integer programming model in the S2 includes: The objective function with the goal of minimizing the time for all UAVs to complete the inspection tasks: where, i and j are task point or virtual site indices, i, j ∈ N ∪ D; N is the task point set; D is the virtual site set, one UAV corresponds to one virtual site, and the virtual site is used to ensure that each UAV starts from the starting point and finally returns to the starting point. The coordinates of the virtual sites are all the starting point coordinates; R ij is the reliable distance between task point i and task point j; F l is the set of task points corresponding to cluster l; m l is the number of task points to be accessed in cluster l; v is the UAV speed; x ij is a decision variable, which takes 1 when task point i is connected to task point j, and 0 otherwise.

3. The multi-UAV three-dimensional path planning method for a task set according to claim 2, wherein The integer programming model in the S2 also includes: Constraint conditions: (1) Each virtual site enters and exits once, that is, each UAV starts from the starting point and then returns to the starting point, (2) If there is an incoming flow at point j, then point j is visited by the UAV. (3) If there is an incoming flow at point i, then point i is also visited by the UAV. (4) A specified number of task points in each cluster need to be visited. (5) Each point can be visited at most once. (6) If the path (i, j) is selected, then points i and j should be visited by the same UAV. (7) Eliminate the sub - tour constraint. where k is the UAV index, k ∈ K; n is the sum of the number of task points and virtual sites; z ik 、z jk respectively represent the order in which task points i and j are visited by drone k; y ik 、y jk are both decision variables. When the drone k visits the task point i, y ik takes 1, otherwise takes 0; when the drone k visits the task point j, y jk takes 1, otherwise takes 0.

4. The multi-UAV three-dimensional path planning method for a task set according to claim 1, characterized in that, In S34, action selection is performed according to the state of the current population, the reward is calculated, and the Q - table is updated. Specifically, it includes: S343. According to the state of the current population, an ε - greedy strategy is used for action selection to determine the crossover probability and mutation probability; the ε - greedy strategy means: where, π(s t , a t ) represents the optimal policy selected in the s t state; Define \(S = \{s_1, \ldots, s_n\}\) as the set of states, where \(n\) is the total number of states; \(A=\{a_1, \ldots, a_m\}\) as the set of candidate actions to be executed, where \(m\) is the total number of actions, and \(a_1=(Pc_1, Pm_1)\), \(a_2=(Pc_2, Pm_2)\), and so on; \(Pc\) is the crossover probability and \(Pm\) is the mutation probability; \(Q(s t , a)\) represents the set of actions that can be selected in state \(s t ; \(\max a (Q(s t , a))\) represents selecting the action with the largest \(Q\)-value in the \(Q\)-table in state \(s t ; \(\xi\) is a random number; \(rand(0, 1)\) represents randomly generating a decimal number between 0 and 1; \(\varepsilon\) represents the exploration rate parameter. S344. Calculate the reward. r = y1 * r1 + y2 * r2 where r is the value of the reward function, y1 and y2 are the weights of the corresponding metric values respectively; r1>0 means that if the objective function value of the optimal individual in the t - th generation is better than that in the (t - 1) - th generation, then the current crossover probability Pc is rewarded; r2>0 means that if the average objective function value in the t - th generation is better than that in the (t - 1) - th generation, then the current mutation probability Pm is rewarded. S345. Update the Q - table. Q(s t ,a t ) = (1 - α)Q(s t ,a t ) + α(r + γmax a Q(s t+1 ,a t+1 )) where Q(s t , a t ) represents the Q-value obtained by taking action a t in state s t , Q(s t+1 , a t+1 ) represents the Q-value obtained by taking action a t+1 in state s t+1 ; α is the learning factor of Q-learning; r is the reward in the current state s t ; γ is the discount factor of the algorithm.

5. A multi-UAV three-dimensional path planning system for a task set, characterized in that, For implementing the multi - UAV three - dimensional path planning method as described in claim 1, it includes: A data acquisition module for acquiring the task set to be inspected and UAV resources. A model construction module for constructing an integer programming model with the goal of minimizing the time for all UAVs to complete the inspection tasks according to the task set to be inspected and UAV resources. A result solving module for solving the integer programming model by using a genetic algorithm integrating the Q - learning mechanism to obtain the multi - UAV three - dimensional path planning result.

6. A storage medium, characterized in that, It stores a computer program for multi - UAV three - dimensional path planning for a task set, where the computer program enables the computer to execute the multi - UAV three - dimensional path planning method for a task set as described in any one of claims 1 to 4.

7. An electronic device, characterized in that, It includes: One or more processors; A memory; And one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include the multi - UAV three - dimensional path planning method for a task set as described in any one of claims 1 to 4.