Air-ground cooperative distribution path planning method based on swarm intelligence optimization

Through a method based on group intelligence optimization, combined with genetic algorithm, clustering and Q-learning algorithm, the joint distribution path of drones and vehicles is optimized, and the problems of long and low efficiency of path planning in the existing technology are solved, and faster and more efficient delivery is achieved.

CN120069248APending Publication Date: 2025-05-30BEIHANG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311601803.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing joint distribution path planning method for drones and vehicles has a long calculation time and low distribution efficiency, making it difficult to find the path with the shortest delivery time in a short time.

Method used

The joint distribution path planning method of air-ground collaborative distribution path based on group intelligence optimization is adopted, and the idea of ​​genetic algorithm is cross-transmitted and mutated, and the combined distribution path of drones and vehicles is optimized.

Benefits of technology

It significantly reduces the calculation time of path planning, improves distribution efficiency, and enables the joint work of drones and vehicles to complete the delivery task in the shortest time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069248A_ABST
    Figure CN120069248A_ABST
Patent Text Reader

Abstract

The invention relates to an air-ground collaborative distribution path planning method based on swarm intelligence optimization, belongs to the technical field of path planning, and solves the problems of long calculation time and low distribution efficiency of an existing unmanned aerial vehicle and vehicle joint distribution path. The method comprises the following steps: minimizing the time of returning an unmanned aerial vehicle or a vehicle to a warehouse at the latest as a target function; randomly generating a plurality of vehicle distribution paths as a parent population; after the parent species are clustered, selecting a parent for crossing to generate a filial population; adopting a Q-learning algorithm to select a neighborhood search operator to mutate the offspring population, and obtaining a mutated offspring population according to the target function; an elite re-insertion strategy is adopted, the parent population is updated according to the filial population, one-time genetic evolution is completed, and the updated parent population is iterated for genetic evolution until the maximum number of iterations is reached; and obtaining an optimal solution in the final parent population according to the target function, and taking the optimal solution as a joint distribution path of the unmanned aerial vehicle and the vehicle. The efficient joint distribution path can be quickly obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of path planning, and particularly to an air-ground collaborative distribution path planning method based on swarm intelligence optimization. Background Art

[0002] The logistics industry in China is developing increasingly, with an unprecedented scale. How to reduce the costs in the logistics transportation process has become an issue that logistics enterprises are closely concerned about. With the rapid development of unmanned aerial vehicle (UAV) technology, the characteristics of UAVs such as flexibility, speed, and low cost have provided huge application space in many industries. Therefore, how to plan an optimal route for the delivery of goods by UAVs in cooperation with vehicles to minimize the total cost of the distribution process is an issue that is currently closely concerned about.

[0003] The existing solutions to the problem of joint distribution path planning for UAVs and vehicles are divided into deterministic methods and heuristic methods. The deterministic methods include non-linear integer programming solvers, dynamic programming, etc.; the heuristic methods include genetic algorithms, greedy algorithms, intelligent optimization methods, etc.

[0004] The advantage of the deterministic method is that the problem to be solved is decomposed into several sub-problems, and the solution of the original problem is obtained by solving the solutions of the sub-problems. Since the solutions of the sub-problems are often not independent of each other, the use of the deterministic method can avoid a large amount of repeated calculations; however, there is no unified processing method for different problems, and it is necessary to analyze and process according to the specific nature of different problems; in addition, when the dimension of the variables increases, the total amount of calculation and storage increases sharply, and these limitations have a large contradiction with the requirements of the joint distribution problem of UAVs and vehicles, and the executability of the planned path is poor, and the distribution efficiency is low.

[0005] Compared with the deterministic method, the heuristic method can iteratively find the optimal path in the current situation by designing a greedy strategy or a heuristic function. However, due to the large number of nodes and high dimensions involved in the joint distribution path planning of UAVs and vehicles, any change in the order of nodes in the distribution path may change the original local optimal path information, and ultimately cause large fluctuations in the convergence time and results, and it is impossible to find the path with the shortest distribution time in a short time. Summary of the Invention

[0006] In view of the above analysis, the embodiments of the present invention aim to provide an air-ground collaborative distribution path planning method based on swarm intelligence optimization to solve the problems of long calculation time and low distribution efficiency in the existing joint distribution of UAVs and vehicles.

[0007] The embodiments of the present invention provide an air-ground collaborative distribution path planning method based on swarm intelligence optimization, including the following steps:

[0008] Taking the minimization of the time when the drone or vehicle returns to the warehouse at the latest as the objective function, a joint distribution mathematical model is constructed;

[0009] Initialize the distribution map and parameters, and randomly generate multiple vehicle distribution paths as the parental population;

[0010] Cluster the parental population, select parents for crossover according to the clustering results to generate the offspring population; based on the mutation probability, use the Q-learning algorithm to select the neighborhood search operator to mutate the offspring population, and obtain the mutated offspring population according to the objective function; adopt the elite reinsertion strategy to update the parental population according to the mutated offspring population, complete one genetic evolution, and iteratively perform genetic evolution on the updated parental population until the maximum number of iterations is reached;

[0011] Obtain the optimal solution in the final parental population according to the objective function as the joint distribution path of the drone and the vehicle.

[0012] Based on the further improvement of the above method, randomly generate multiple vehicle distribution paths, including:

[0013] Based on the coordinates of the warehouse and each customer point in the distribution map, and the vehicle speed in the parameters, according to the time when the vehicle departs from the warehouse, passes through each customer point and returns to the warehouse, generate the path with the shortest time as the optimal vehicle distribution path, and then randomly shuffle the order of each customer point on the path to generate the most non - identical vehicle distribution paths.

[0014] Based on the further improvement of the above method, cluster the parental population, including: merge the vehicle distribution path and the drone distribution path of each parent in the parental population, and based on the K - medoids clustering algorithm, calculate the distance from each parent to the class center according to the Levenshtein distance algorithm, and perform clustering division on the parental population.

[0015] Based on the further improvement of the above method, select parents for crossover according to the clustering results to generate the offspring population, including:

[0016] Select two parents that do not belong to the same cluster. After merging the vehicle distribution paths and drone distribution paths of the two selected parents respectively, use the partial crossover matching operator for crossover. Taking one parent as the first parent, when there is no crossover conflict, obtain the initial path; if the first parent does not have a drone distribution path, the initial path is put into the offspring population as the offspring. If there is, according to the number of drone distribution paths of the first parent, obtain the drone distribution path and the vehicle distribution path from the initial path and put them into the offspring population as the offspring.

[0017] Repeat the above steps until the number of the offspring population is the same as that of the parental population.

[0018] Based on further improvements to the above method, according to the number of drone delivery routes of the first parent generation, obtain drone delivery routes and vehicle delivery routes from the initial routes, including:

[0019] According to the number of drone delivery routes, in accordance with the merging rules of vehicle delivery routes and drone delivery routes, screen out the customer points in the drone delivery routes from the initial routes, and obtain vehicle delivery routes based on the remaining nodes of the initial routes; use the objective function as the selection function of the greedy algorithm, and select take-off nodes and landing nodes for each customer point from the nodes of the vehicle delivery routes to obtain drone delivery routes.

[0020] Based on further improvements to the above method, use the Q-learning algorithm to select a neighborhood search operator to mutate the offspring population, and obtain the mutated offspring population according to the objective function, including:

[0021] Based on the global Q-table, according to the neighborhood search operator of the current state, select the action of the current state according to the greedy strategy, use the neighborhood search operator corresponding to the action as the neighborhood search operator of the next state, update the global Q-table, and select the neighborhood search operator of the next state to search for the offspring, generate multiple neighborhood individuals, and take the neighborhood individual with the smallest objective function value as the mutated offspring to update the offspring population.

[0022] Based on further improvements to the above method, the global Q-table is a square matrix constructed with neighborhood search operators as states and actions, and the Q values in the global Q-table are initially random numbers in [0, 0.1].

[0023] Based on further improvements to the above method, updating the global Q-table includes: searching for the offspring according to the neighborhood search operator of the current state and the neighborhood search operator of the next state respectively, calculating the minimum objective function values of the two states, and calculating the reward value according to the minimum objective function values of the two states; updating the Q value of the current state in the global Q-table according to the Q value of the current state, the maximum Q value of the next state, and the reward value.

[0024] Based on further improvements to the above method, according to the minimum objective function values of the two states, calculate the reward value through the following formula:

[0025]

[0026] where r represents the reward value, β represents a preset constant, t c+1 represents the minimum objective function value of the current state, represents the minimum objective function value of the next state.

[0027] Based on further improvements to the above method, adopt the elite reinsertion strategy to update the parent population according to the mutated offspring population, including:

[0028] Based on the objective function, sort the parent population and the offspring population in ascending order of the objective function value, and select the first half of the individuals in the offspring population to replace the second half of the individuals in the parent population.

[0029] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects: using the idea of genetic algorithm, performing crossover and mutation in the genetic process on the initial distribution path with only vehicles, and randomly restoring it to the combined path of drones and vehicles. Among them, the clustering idea is applied, and the Levenshtein distance is used as the similarity function to prevent the crossover of parents with high path similarity, and the memetic algorithm is adopted to improve the diversity and feasibility of solutions through the crossover operator; using Q-learning to guide the selection of targeted mutation operators for neighborhood search, reducing the fluctuations in the search process, accelerating the convergence speed of solutions, and preventing falling into local optima; selecting the optimal solution starting from the time consumed when drones and vehicles work together to maximize the working efficiency of drones cooperating with vehicles.

[0030] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can be made obvious from the description, or understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained from the content specifically pointed out in the description and the drawings. Description of the Drawings

[0031] The drawings are only for the purpose of showing specific embodiments and are not considered as a limitation of the present invention. Throughout the drawings, the same reference signs represent the same components.

[0032] Figure 1 It is a flowchart of a method for planning an air-ground collaborative distribution path based on swarm intelligence optimization in an embodiment of the present invention;

[0033] Figure 2 It is a schematic diagram of the combined distribution path of drones and vehicles in the distribution map in an embodiment of the present invention. Detailed Embodiments

[0034] The following will specifically describe the preferred embodiments of the present invention with reference to the drawings. Among them, the drawings form a part of this application and are used together with the embodiments of the present invention to explain the principle of the present invention, and are not used to limit the scope of the present invention.

[0035] A specific embodiment of the present invention discloses a method for planning an air-ground collaborative distribution path based on swarm intelligence optimization, as Figure 1 shown, including the following steps:

[0036] S1. Construct a joint distribution mathematical model with the objective of minimizing the time when the drone or vehicle returns to the warehouse at the latest.

[0037] It should be noted that in this embodiment, the collaborative distribution between the open space and the vehicle mainly refers to the joint distribution of the drone and the vehicle. The drone or vehicle departs from the warehouse. Under the condition that each customer point is visited by the vehicle or drone and only served once, according to the planned vehicle distribution path and drone distribution path, the total time to finally return to the warehouse is the shortest; among them, the drone takes off from the vehicle or warehouse, serves only one customer, and returns to the vehicle or warehouse; the vehicle is responsible for distribution, charging the drone and providing goods, and the drone only participates in distribution. Therefore, in the constructed single-vehicle and single-drone joint transportation and distribution mathematical model, with the objective of minimizing the time when the drone or vehicle returns to the warehouse at the latest, it is expressed as follows:

[0038] Min t c+1 Formula (1),

[0039] where, t c+1 represents the time when the drone or vehicle returns to the warehouse at the latest, c represents the number of customer points, and c + 1 represents the drone or vehicle returning to the warehouse.

[0040] In this embodiment, C is used to represent the set of all customer points {1, 2,..., c}, and C' is used to represent the set of customer points served by the drone The vehicle / drone departing from the warehouse is represented by 0, and returning to the warehouse is represented by c + 1. N represents the set of all nodes {0, 1, 2,..., c + 1}, N 0 represents the set of legal departure / takeoff nodes {0, 1, 2,..., c}; N + represents the set of legal landing nodes of the drone {1, 2,..., c + 1}; P represents the set of nodes corresponding to the drone distribution path (i, j, k), where i represents the takeoff node, j represents the service node, and k represents the landing node; u i represents the absolute position of node i in the vehicle distribution path, 1 ≤ u i ≤ c + 2, t i represents the total time when the vehicle arrives at node i, t i ≥ 0, t' i represents the total time when the drone arrives at node i, t' i ≥ 0, τ represents the vehicle speed, τ’ represents the drone speed, S L represents the time taken for the vehicle to launch the drone, S R represents the time taken for the vehicle to receive the drone, M ∞Represents an "infinity" number, which is used in the time-related constraint conditions in this embodiment. Exemplarily, it is preset to 3000 seconds.

[0041] When calculating the objective function value, the following constraint conditions are established according to the application scenario:

[0042] (1) Each customer can only be served once, which is expressed as follows:

[0043]

[0044] Among them, x ij Indicates whether the vehicle departs from node i to serve node j. If it serves, take 1; otherwise, take 0, that is: x ij ∈{0,1}, j∈{N + :j≠i}; y ijk Indicates whether the drone departs from node i to serve node j and returns to node k. If it serves, take 1; otherwise, take 0, that is: y ijk ∈{0,1}, j∈{C:j≠i},k∈{M + :(i,j,k)∈P};

[0045] (2) The vehicle only departs from the warehouse once, which is expressed as follows:

[0046]

[0047] (3) The vehicle only returns to the warehouse once, which is expressed as follows:

[0048]

[0049] (4) To eliminate sub-loops, if the vehicle goes from node i to node j, then the absolute position of node i in the vehicle delivery path must be one before node j, which is expressed as follows:

[0050]

[0051] (5) The vehicle serving a node must leave that node, that is, the in-degree of the vehicle at the customer point is equal to the out-degree, which is expressed as follows:

[0052]

[0053] (6) The drone can be launched at most once from any specific node (including the warehouse), which is expressed as follows:

[0054]

[0055] (7) The drone can rendezvous with any specific node (including the customer and the end warehouse) at most once. That is, for the same customer point / end warehouse, the drone can visit it at most once, as shown below:

[0056]

[0057] (8) If the drone departs from customer point i and is received by the vehicle at node k, then nodes i and k must be assigned to the vehicle simultaneously, as shown below:

[0058]

[0059] (9) If the drone departs from the warehouse to serve node j and lands at node k, it must be received by the vehicle at node k, as shown below:

[0060]

[0061] (10) If the drone takes off from node i and is received by the vehicle at node k, then the vehicle must serve node i before serving node k, as shown below:

[0062]

[0063] (11) Synchronization constraint between the drone's takeoff point and the vehicle: If the drone takes off from node i, then the vehicle and the drone arrive at node i at the same time, as shown below:

[0064]

[0065] (12) Synchronization constraint between the drone's landing point and the vehicle's receiving point: If the drone lands at node k, then the vehicle and the drone arrive at node k at the same time, as shown below:

[0066]

[0067] (13) The vehicle's effective arrival time at customer point k: If the vehicle travels from node h to node k, then the arrival time at node k is the departure time from point h plus the travel time between h and k. If the drone takes off or lands at node k, it also needs to include the corresponding vehicle launching drone time or receiving drone time, as shown below:

[0068]

[0069] (14) If the drone takes off from node i, then the arrival time of the drone at node j must include the flight time from node i to node j, as shown below:

[0070]

[0071] (15) If the UAV is received by the vehicle at node k, then the arrival time at node k must include the travel time from node j to node k plus the reception time at node k, as shown below:

[0072]

[0073] (16) UAV endurance constraint, ensuring that the flight time of the UAV does not exceed its maximum service time e, as shown below:

[0074]

[0075] (17) Order constraint between nodes i and j in the vehicle delivery route, as shown below:

[0076]

[0077] where p ij represents the access order in the vehicle delivery route between nodes i and j. If node i is before node j, take 1; otherwise, take 0, that is: p ij ∈{0,1}, j∈{C:j≠i}; p 0j =1,

[0078] (18) UAV launch time constraint, ensuring that the landing time of the UAV is later than the takeoff time of the UAV, as shown below:

[0079]

[0080] (19) The departure time t 0 of the vehicle from the warehouse is 0, and the departure time t' 0 of the UAV from the warehouse is 0, as shown below:

[0081] t 0 =0, t' 0 =0 Formula (20).

[0082] S2. Initialize the delivery map and parameters, and randomly generate multiple vehicle delivery routes as the parent population.

[0083] It should be noted that the delivery map includes 1 warehouse and c customer points, as well as the routes between the warehouse and the customer points. Initializing the delivery map includes: taking the warehouse as the origin of the coordinate system, initializing the coordinates of each customer point, establishing the edges between nodes according to the routes, and storing them as an undirected graph.

[0084] The initialization parameters are the parameters for initializing the vehicle and the drone, as well as the parameters in the hybrid genetic algorithm, including: vehicle speed τ, drone speed τ', time S for the vehicle to send the drone L 、time S for the vehicle to receive the drone R 、maximum number of iterations D, population size M, mutation probability P m 、learning rate α, reward decay coefficient γ, greedy coefficient ε and global Q-table in the Q-learning method.

[0085] Next, multiple vehicle delivery routes are randomly generated, including:

[0086] Based on the coordinates of the warehouse and each customer point in the delivery map, and the vehicle speed in the parameters, according to the time for the vehicle to start from the warehouse, pass through each customer point and return to the warehouse, the path with the shortest time is generated as the optimal vehicle delivery route. Then, according to the optimal vehicle delivery route, the order of each customer point on the path is randomly shuffled to generate the maximum number of distinct vehicle delivery routes.

[0087] Exemplarily, the mathematical programming solver Gurobi is used to generate the optimal vehicle delivery route.

[0088] It should be noted that the maximum number of distinct vehicle delivery routes randomly generated initially is the population size M, and the initial drone delivery routes are all empty. In step S3, as the genetic evolution progresses, the vehicle delivery routes are gradually converted into combined vehicle and drone delivery routes.

[0089] Specifically, the combined delivery route includes a vehicle delivery route and a drone speed route. Among them, the starting and ending nodes of the vehicle delivery route are the warehouse, and the intermediate nodes are the customer points served by the vehicle; the drone delivery route is in a triple format, successively including a takeoff node, a service node, and a landing node; the takeoff node and the landing node are the warehouse or a customer point, and the service node is a customer point.

[0090] Exemplarily, in Figure 2 the delivery map, the square Depot represents the warehouse, and 0 is used as the identifier of the warehouse; the circles represent customer points, and 1-9 are used as the identifiers of the customer points respectively. The solid line represents the vehicle route, starting from the warehouse and finally returning to the warehouse, and the dashed line represents the drone route. Then there is 1 vehicle delivery route, which is expressed as: (0,1,2,3,5,6,7,9,0); there are 2 drone delivery routes. One takes off from node 3, serves node 4 and lands at node 5, which is expressed as (3,4,5), and the other takes off from node 6, serves node 8 and lands at node 9, which is expressed as (6,8,9).

[0091] S3. Cluster the parental population, select parents for crossover according to the clustering results to generate an offspring population; based on the mutation probability, use the Q-learning algorithm to select a neighborhood search operator to mutate the offspring population, and obtain the mutated offspring population according to the objective function; adopt an elite reinsertion strategy to update the parental population according to the mutated offspring population to complete one genetic evolution; iteratively perform genetic evolution on the updated parental population until the maximum number of iterations is reached.

[0092] It should be noted that step S3 is an iterative process. Each iteration performs genetic evolution on the currently latest parental population, which sequentially includes four steps of processing: step S31 clustering, step S32 crossover, step S33 mutation, and step S34 update. The following will be described in detail respectively.

[0093] S31. Cluster the parental population.

[0094] It should be noted that clustering the parental population includes: merging the vehicle delivery routes and the drone delivery routes of each parent in the parental population, and based on the K-medoids clustering algorithm, calculating the distance from each parent to the class center according to the Levenshtein distance algorithm, and performing clustering division on the parental population.

[0095] Specifically, merging the vehicle delivery routes and the drone delivery routes of each parent in the parental population means removing the warehouse node and the takeoff and landing nodes of the drone and then merging the remaining customer points, that is, sequentially merging the customer points served by the vehicle and the drone.

[0096] Preferably, after removing the start and end nodes in the vehicle delivery route for each parent, splice the service nodes in the drone delivery route behind. For Figure 2 example, the merged route is (1, 2, 3, 5, 6, 7, 9, 4, 8).

[0097] Based on the K-medoids clustering algorithm, calculating the distance from each parent to the class center according to the Levenshtein distance algorithm and performing clustering division on the parental population includes:

[0098] ① Randomly select parents as the initial class centers of each cluster according to the number of clusters; exemplarily, the number of clusters is taken as M / 5, where M represents the population size.

[0099] ② According to the Levenshtein distance algorithm Levenshtein, calculate the distances from other parents to each class center, and assign them to the class closest to the class center.

[0100] ③ In each class, calculate the total distance between each parent and the remaining parents respectively, and select the parent with the smallest total distance as the new class center; for example, if there are 3 parents in a class, and the distances between the parents form a 3×3 matrix, then the parent with the smallest sum of row values is the class center of this class.

[0101] ④ Repeat steps ② and ③ until the class centers of all classes no longer change or the set maximum number of iterations is reached.

[0102] It should be noted that considering that the objective function values of the combined distribution paths of some drones and vehicles may vary greatly, but the paths may not differ much. If the objective function value is used as the clustering condition, it is not conducive to generating new offspring. Therefore, in this step, the Levenshtein distance algorithm is used for clustering, and it is used as the similarity evaluation function. Parents in the same class are considered to have high similarity.

[0103] S32. Select parents according to the clustering results for crossover to generate an offspring population.

[0104] It should be noted that in this embodiment, the size of the offspring population is the same as that of the parent population. Therefore, in this step, according to the clustering results of the parent population, the parents are repeatedly selected for crossover operations until the number of the offspring population is the same as that of the parent population, that is, M offspring are generated.

[0105] Specifically, select two parents that do not belong to the same cluster according to the clustering results. Preferably, the roulette wheel method is used to select two parents. If the two parents belong to the same cluster, reselect until the two parents do not belong to the same cluster. Prevent parents with high path similarity from performing crossover.

[0106] After merging the vehicle distribution paths and drone distribution paths of the two selected parents respectively, use the partially - matched crossover (PMX) operator for crossover. Take one of the parents as the first parent. When there is no crossover conflict, obtain the initial path; if the first parent does not have a drone distribution path, the initial path is put into the offspring population as an offspring. If there is, according to the number of drone distribution paths of the first parent, obtain the drone distribution path and vehicle distribution path from the initial path and put them into the offspring population as an offspring. Repeat selecting parents for crossover until the number of the offspring population is the same as that of the parent population.

[0107] It should be noted that after merging the vehicle distribution paths and drone distribution paths of each parent, randomly select two crossover points in the two parents I 1 and I 2 to determine the matching segments, and define the mapping relationship between nodes according to the selected matching segments. First, exchange the matching segments of the two parents, and then for parent I1 Conflict detection is performed on other node bits outside the matching segment. If there is a conflict, that is, there are identical nodes, then according to the mapping relationship, the nodes on the corresponding bits are obtained through exchanges in sequence, so that the parent generation I 1 There are no identical nodes, and the initial path is obtained.

[0108] Exemplarily, the parent generation I 1 The vehicle delivery path is (0, 1, 2, 3, 4, 5, 6, 7, 0), the drone delivery path is (1, 8, 4) and (5, 9, 7), and the combined path is (1, 2, 3, 4, 5, 6, 7, 8, 9); the parent generation I 2 The vehicle delivery path is (0, 3, 5, 8, 1, 7, 4, 2, 6, 0), the drone delivery path is (3, 9, 8), and the combined path is (3, 5, 8, 1, 7, 4, 2, 6, 9); the paths after combining the two parent generations correspond one by one according to the node bits. Select 2 and 5 in the parent generation I 1 as the crossover points, then according to the matching segment (2, 3, 4, 5) in the parent generation I 1 and the matching segment (5, 8, 1, 7) in the parent generation I 2 to establish the mapping relationship between nodes. Select the parent generation I 1 as the first parent generation, exchange the two matching segments, and the first parent generation becomes (1, 5, 8, 1, 7, 6, 7, 8, 9), there are conflicts (two 1s, 7s, and 8s). According to the mapping relationship, convert the 1 outside the matching segment to 4, 7 to 2, and 8 to 3, and obtain the initial path (4, 5, 8, 1, 7, 6, 2, 3, 9) without conflicts.

[0109] Next, if the first parent generation does not have a drone delivery path, the initial path is put into the offspring population as an offspring. Since the warehouse nodes of the vehicle delivery path are removed when merging the paths, at this time, the warehouse nodes can be added before and after the initial path as the starting and ending nodes and put into the offspring population as an offspring.

[0110] If the first parent generation has a drone delivery path, then according to the number of drone delivery paths of the first parent generation, obtain the drone delivery path and the vehicle delivery path from the initial path, including:

[0111] According to the number of drone delivery paths, in accordance with the merging rules of the vehicle delivery path and the drone delivery path, screen out the customer points in the drone delivery path from the initial path, and obtain the vehicle delivery path according to the remaining nodes of the initial path; use the objective function as the selection function of the greedy algorithm to select the takeoff node and the landing node for each customer point from the nodes of the vehicle delivery path to obtain the drone delivery path.

[0112] Exemplarily, in the above example, the first parent generation I1 There are 2 UAV delivery paths. When merging, the customer points served by the UAVs are merged behind the customer points served by the vehicle. Therefore, the last two nodes 3 and 9 of the initial path are selected as the customer points in the UAV delivery path. According to the remaining (4, 5, 8, 1, 7, 6, 2), the vehicle delivery path (0, 4, 5, 8, 1, 7, 6, 2, 0) can be obtained; then, using the objective function as the selection function of the greedy algorithm, the take-off node and the landing node are selected from (4, 5, 8, 1, 7, 6, 2) for nodes 3 and 9 respectively, and the UAV delivery paths are obtained, such as (4, 3, 8) and (1, 9, 7). The vehicle delivery path obtained from the initial path and the 2 UAV delivery paths are used as the generated offspring and put into the offspring population.

[0113] S33. Based on the mutation probability, the Q-learning algorithm is used to select the neighborhood search operator to mutate the offspring population, and the mutated offspring population is obtained according to the objective function.

[0114] It should be noted that in this embodiment, there are 9 neighborhood search operators, and mutating the offspring is to perform neighborhood search on the offspring according to the selected neighborhood search operator.

[0115] Specifically, the 9 neighborhood search operators include:

[0116] Ω 1 Vehicle node swapping operator: Swap any two selected vehicle nodes;

[0117] Ω 2 Vehicle node insertion operator: Specify a node and two adjacent nodes, and insert the selected node between the selected two adjacent nodes;

[0118] Ω 3 Two-node optimization operator: Reverse all the nodes between any two selected vehicle nodes;

[0119] Ω 4 UAV node removal operator: Remove the selected UAV node and randomly add it to the vehicle path;

[0120] Ω 5 UAV node addition operator: Remove any one vehicle node, use it as the service node of the UAV, and use the greedy algorithm to allocate the take-off node and the landing node for it to obtain the UAV triple;

[0121] Ω 6 UAV node swapping operator: Swap the service nodes in any two selected UAV triples;

[0122] Ω 7UAV takeoff and landing node allocation: Randomly select a UAV triple, and reselect the takeoff node and landing node in the specified triple;

[0123] Ω 8 Vehicle consecutive three-node and UAV triple exchange operator: Exchange any three consecutive vehicle nodes that are not triples, randomly select a UAV triple, and exchange them in the selected order;

[0124] Ω 9 UAV takeoff node forward movement or landing node backward movement operator: Randomly select a UAV triple, reselect the takeoff node or landing node in the specified triple, and move the UAV takeoff node forward or move the landing node backward.

[0125] It should be noted that in this embodiment, a Q-learning algorithm is used to guide the selection of a targeted neighborhood search operator for neighborhood search instead of randomly selecting a neighborhood search operator, reducing the fluctuations in the search process, accelerating the convergence speed of the solution, and preventing falling into a local optimum.

[0126] In this embodiment, the neighborhood search operator is used as the state and action in the Q-learning algorithm to establish a global Q-table. Therefore, the Q-table is a 9×9 square matrix, and the Q values in the global Q-table are initially random numbers in [0, 0.1], representing the value that can be brought by selecting each action in each state.

[0127] Generate a probability value randomly for each offspring. If it is greater than the mutation probability, the offspring will not mutate; otherwise, it will mutate. Among them, for the first offspring to mutate, randomly select a neighborhood search operator as the current state. For the subsequent offspring to mutate, including the subsequent genetic evolution process, the neighborhood search operator selected by the previous offspring is used as the current state.

[0128] The mutation operation includes: Based on the global Q-table, according to the neighborhood search operator in the current state, select the action in the current state according to the greedy strategy. The neighborhood search operator corresponding to this action is used as the neighborhood search operator in the next state. Update the global Q-table, and select the neighborhood search operator in the next state to search for the offspring, generating multiple neighborhood individuals. Select the neighborhood individual with the smallest objective function value as the mutated offspring, and update the offspring population.

[0129] It should be noted that to prevent falling into a local optimum, the parameter ε is set for the greedy strategy selection. That is, there is a probability of ε for random selection and a probability of 1 - ε for selection according to the Q-table. Exemplarily, when the parameter ε is set to 10%, when the randomly generated probability is in [0, 10%], randomly select an action (neighborhood search operator), and when the randomly generated probability is in (10%, 100%], select the action (neighborhood search operator) with the largest Q value corresponding to this state from the Q-table.

[0130] Furthermore, updating the global Q-table includes: searching for offspring according to the neighborhood search operator of the current state and the neighborhood search operator of the next state respectively, calculating the minimum objective function values of the two states, and calculating the reward value through formula (21) according to the minimum objective function values of the two states; updating the Q-value of the current state in the global Q-table through formula (22) according to the Q-value of the current state, the maximum Q-value of the next state, and the reward value.

[0131]

[0132] Among them, r represents the reward value, β represents a preset constant, t c+1 represents the minimum objective function value of the current state, represents the minimum objective function value of the next state. That is, when the next state takes less time, it is a positive reward, the Q-value will increase, and the probability of being selected will increase; on the contrary, it is a negative reward, the Q-value will decrease, but the corresponding neighborhood search operator is still selected to perform neighborhood search on the offspring.

[0133] Q(s,a) = Q(s,a) + α * (r + γ * max a' (Q(s',a')) - Q(s,a)) Formula (22),

[0134] Among them, s represents the current state, a represents the current action, s' and a' represent the next state and action respectively, α represents the learning rate, γ represents the reward decay coefficient, Q(s,a) represents the Q-value of the current state, and max a' (Q(s',a')) represents the maximum Q-value of the next state.

[0135] S34. Adopt the elite reinsertion strategy to update the parent population according to the mutated offspring population.

[0136] Specifically, based on the objective function, sort the parent population and the offspring population from small to large according to the objective function value. The smaller the objective function value, the better the individual. Take the first half of the individuals in the offspring population to replace the second half of the individuals in the parent population, and update the parent population, that is, complete one genetic evolution.

[0137] Repeat step S3 for the updated parent population and update it again until the maximum number of iterations is reached.

[0138] S4. Obtain the optimal solution in the final parent population according to the objective function as the combined distribution path of the unmanned aerial vehicle and the vehicle.

[0139] It should be noted that after the iteration of step S3 ends, the finally updated parent population is used as the final parent population. The objective function value of each individual in it is calculated, and the individual with the smallest objective function value is taken as the optimal solution to obtain the combined distribution path of the unmanned aerial vehicle and the vehicle.

[0140] Compared with the prior art, the method for planning the collaborative distribution path between air and ground based on swarm intelligence optimization provided in this embodiment uses the idea of genetic algorithm to perform crossover and mutation in the genetic process on the initial distribution path with only vehicles, and randomly restores it to the combined path of the unmanned aerial vehicle and the vehicle. Among them, the clustering idea is applied, and the Levenshtein distance is used as the similarity function to prevent the crossover of parents with high path similarity, and the memetic algorithm is adopted to improve the diversity and feasibility of the solution through the crossover operator; Q-learning is used to guide the selection of targeted mutation operators for neighborhood search, reduce the fluctuations in the search process, accelerate the convergence speed of the solution, and prevent falling into local optimum; the optimal solution is selected starting from the time consumed when the unmanned aerial vehicle and the vehicle work together to maximize the working efficiency of the unmanned aerial vehicle cooperating with the vehicle.

[0141] Those skilled in the art can understand that all or part of the processes for implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory or a random access memory, etc.

[0142] The above is only a specific and preferred embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.

Claims

1. An air-ground collaborative distribution path planning method based on swarm intelligence optimization, characterized in that, it includes the following steps: Construct a joint distribution mathematical model with the objective of minimizing the time when the drone or vehicle returns to the warehouse at the latest; Initialize the distribution map and parameters, and randomly generate multiple vehicle distribution paths as the parent population; Cluster the parent population, select parents for crossover according to the clustering results to generate the offspring population; based on the mutation probability, use the Q-learning algorithm to select the neighborhood search operator to mutate the offspring population, and obtain the mutated offspring population according to the objective function; adopt the elite reinsertion strategy to update the parent population according to the mutated offspring population to complete one genetic evolution; iterate the updated parent population for genetic evolution until the maximum number of iterations is reached; Obtain the optimal solution in the final parent population according to the objective function as the joint distribution path of the drone and the vehicle.

2. The air-ground collaborative distribution path planning method based on swarm intelligence optimization according to claim 1, characterized in that, the random generation of multiple vehicle distribution paths includes: Based on the coordinates of the warehouse and each customer point in the distribution map, and the vehicle speed in the parameters, generate the path with the shortest time as the optimal vehicle distribution path according to the time when the vehicle departs from the warehouse, passes through each customer point and returns to the warehouse, and then randomly shuffle the order of each customer point on the optimal vehicle distribution path to generate the maximum number of mutually different vehicle distribution paths.

3. The air-ground collaborative distribution path planning method based on swarm intelligence optimization according to claim 1, characterized in that, the clustering of the parent population includes: merging the vehicle distribution path and the drone distribution path of each parent in the parent population, and based on the K-medoids clustering algorithm, calculating the distance from each parent to the class center according to the Levenshtein distance algorithm, and clustering and dividing the parent population.

4. The air-ground collaborative distribution path planning method based on swarm intelligence optimization according to claim 1, characterized in that, the selection of parents for crossover according to the clustering results to generate the offspring population includes: Select two parents that do not belong to the same cluster. After merging the vehicle distribution paths and drone distribution paths of the two selected parents respectively, use the partial crossover matching operator for crossover. Take one of the parents as the first parent. When there is no crossover conflict, obtain the initial path; if the first parent does not have a drone distribution path, the initial path is put into the offspring population as the offspring. If there is, according to the number of drone distribution paths of the first parent, obtain the drone distribution path and the vehicle distribution path from the initial path and put them into the offspring population as the offspring; Repeat the above steps until the number of the offspring population is the same as the number of the parent population.

5. The air-ground collaborative distribution path planning method based on swarm intelligence optimization according to claim 4, characterized in that, the obtaining of the drone distribution path and the vehicle distribution path from the initial path according to the number of drone distribution paths of the first parent includes: According to the number of UAV delivery routes, in accordance with the merging rules of vehicle delivery routes and UAV delivery routes, screen out the customer points in the UAV delivery routes from the initial routes, and obtain the vehicle delivery routes based on the remaining nodes of the initial routes; use the objective function as the selection function of the greedy algorithm to select the take-off node and landing node for each customer point from the nodes of the vehicle delivery routes to obtain the UAV delivery routes.

6. The method for collaborative air-ground delivery route planning based on swarm intelligence optimization according to claim 1, characterized in that the step of using the Q-learning algorithm to select a neighborhood search operator to mutate the offspring population and obtaining the mutated offspring population according to the objective function includes: Based on the global Q-table, according to the neighborhood search operator of the current state, select the action of the current state according to the greedy strategy, and use the neighborhood search operator corresponding to the action as the neighborhood search operator of the next state, update the global Q-table, and select the neighborhood search operator of the next state to search for the offspring, generate multiple neighborhood individuals, and take the neighborhood individual with the smallest objective function value as the mutated offspring to update the offspring population.

7. The method for collaborative air-ground delivery route planning based on swarm intelligence optimization according to claim 6, characterized in that the global Q-table is a square matrix constructed with neighborhood search operators as states and actions, and the Q values in the global Q-table are initially random numbers in [0, 0.1].

8. The method for collaborative air-ground delivery route planning based on swarm intelligence optimization according to claim 6, characterized in that the step of updating the global Q-table includes: searching for the offspring according to the neighborhood search operator of the current state and the neighborhood search operator of the next state respectively, calculating the minimum objective function values of the two states, calculating the reward value according to the minimum objective function values of the two states; updating the Q value of the current state in the global Q-table according to the Q value of the current state, the maximum Q value of the next state and the reward value.

9. The method for collaborative air-ground delivery route planning based on swarm intelligence optimization according to claim 8, characterized in that the reward value is calculated by the following formula according to the minimum objective function values of the two states: where r represents the reward value, β represents a preset constant, and t c+1 represents the minimum objective function value of the current state, and represents the minimum objective function value of the next state.

10. The method for collaborative air-ground delivery route planning based on swarm intelligence optimization according to claim 1, characterized in that the step of using the elite reinsertion strategy to update the parent population according to the mutated offspring population includes: Based on the objective function, sort the parent population and the offspring population from small to large according to the objective function values, and take the first half of the individuals in the offspring population to replace the second half of the individuals in the parent population.

Citation Information

Cited By

  • Multi-unmanned aerial vehicle cooperative task allocation method, device, equipment and medium

    CN121168978A

  • Logistics distribution path optimization method based on double-improved genetic simulated annealing algorithm

    CN122198814A