Vehicle-unmanned aerial vehicle logistics distribution path planning method based on deep reinforcement learning
Through deep reinforcement learning and Actor-Critic algorithm, and combining with the K-means clustering method to allocate tasks, the existing path planning methods in terms of computational efficiency and path quality are solved, and efficient and accurate distribution path optimization is achieved.
Patent Information
- Application Number
- CN202510112625.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing vehicle-drone combined distribution path planning method has shortcomings in computing efficiency and path quality, especially in large-scale distribution tasks and complex scenarios.
The automobile-drone path planning method based on deep reinforcement learning is adopted. By building a delivery scenario, designing a reinforcement learning model, and using the Actor-Critic method for optimization training, the drone route and vehicle path are optimized, and task allocation is combined with the K-means clustering method.
It realizes efficient and accurate distribution path optimization, improves the overall efficiency and energy efficiency of the logistics distribution system, and can flexibly respond to dynamic changes in complex scenarios.
Smart Images

Figure CN119990493A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning technology, and relates to a vehicle-UAV path planning method based on deep reinforcement learning, and in particular to a vehicle and UAV collaborative delivery path planning method in logistics delivery. Background Art
[0002] With the rapid development of e-commerce, the logistics industry has a broad market demand, especially the increasing requirements for efficiency and flexibility in the "last mile" delivery. Traditional ground transportation modes often have great limitations when facing complex traffic environments and efficient and timely delivery needs. Especially in urban delivery tasks, factors such as traffic congestion and limited delivery route selection have greatly reduced delivery efficiency and cost-effectiveness. Drone delivery, with its flexibility and efficiency, can bypass the limitations of ground transportation and greatly improve the flexibility and timeliness of delivery. Especially in remote areas or urban terminal delivery, drones can not only reach customer locations quickly, but also significantly reduce delivery costs. However, the limited load and battery life of drones make it difficult for drones to independently undertake all work in large-scale delivery tasks. Therefore, the combination of car and drone delivery mode has become an innovative solution with great potential. This mode combines the efficiency of ground transportation tools and the flexibility of drones, and can give full play to their respective advantages in different delivery tasks. Vehicles can undertake the take-off and landing and charging tasks of drones, and carry drones to complete specific delivery needs.
[0003] However, vehicle-mounted drone delivery faces challenges in path planning. First, the problem of collaborative optimization between vehicles and drones, especially in path planning, how to reasonably allocate delivery tasks to customer nodes to avoid duplicate transportation or uneven distribution, is a difficult problem that needs to be solved urgently. Second, most existing path planning methods are based on precise algorithms or heuristic algorithms, but when faced with large-scale delivery tasks, these methods usually have low computational efficiency or poor solution quality and cannot effectively cope with the delivery needs of complex scenarios.
[0004] In view of the current status of research and development in this field, a reinforcement learning-based car-UAV logistics distribution path planning method is designed. A single distribution method is easily limited by road resources, resulting in traffic congestion and low distribution efficiency. The distribution method combining cars and drones can overcome the limitations of a single method, reduce the dependence of vehicles on road resources, reduce traffic congestion, and provide a more flexible and economical distribution solution. At the same time, in the planning of distribution paths, traditional path planning methods often face problems such as high computational complexity, difficulty in handling dynamic changes, and difficulty in balancing the coordinated optimization of vehicles and drones in complex distribution environments. Reinforcement learning methods can overcome the limitations of traditional path planning methods in dealing with complex distribution environments, and can continuously optimize decision-making strategies through interaction with the environment. This method makes full use of the learning ability of reinforcement learning agents, combines the characteristics of cars and drones, and can achieve distribution path optimization and energy consumption reduction under the constraints of drone load and energy consumption. Therefore, this method provides an innovative and efficient intelligent solution for car-UAV collaborative distribution in the logistics industry, which has great research significance and application prospects. Summary of the invention
[0005] In order to solve the problems of low computational efficiency, poor path quality and inability to adapt to complex demand scenarios faced by the existing traditional path planning method for combined car-drone delivery, the present invention proposes a car-drone combined path planning optimization method based on deep reinforcement learning. The method first constructs a delivery scenario, generates a data set of the delivery environment, and performs task allocation to obtain the car's stop points and the delivery tasks of the drone; then, a reinforcement learning model is designed for the delivery route planning task of the drone and the path planning task of the car between the stop points, and the Actor-Critic method is used for optimization training; the drone model obtains the optimal route planning strategy through learning and training, and then uses the drone route planning strategy to integrate into the car path planning training, and finally obtains the collaborative delivery method of the car and the drone. The present invention combines K-means clustering and Actor-Critic algorithm, optimizes the path planning strategy by training the intelligent agent, and considers the energy, load and other constraints of the car and the drone at the same time, which can efficiently and accurately optimize the delivery path and improve the overall efficiency and energy efficiency of the logistics distribution system.
[0006] The technical solution implemented by the present invention is as follows:
[0007] A method for car-UAV logistics distribution path planning based on deep reinforcement learning, comprising:
[0008] S1. Construct a logistics delivery scenario of a car-drone combination, i.e. generate a delivery environment data set; assign tasks to cars and drones based on the location of customer nodes, i.e. cluster the customer nodes in the delivery environment data set using the K-means clustering method, with each cluster center serving as the location of the car stop, and each clustering cluster serving as the delivery area of the drone at the corresponding car stop;
[0009] S2. For the UAV route planning task between the car stop and the customer node, a reinforcement learning model for UAV route planning is constructed by designing the state space, action space and reward function;
[0010] S3. Based on the UAV route planning reinforcement learning model, the policy network and the value network are optimized by the Actor-Critic method to train and optimize the UAV route planning policy model;
[0011] S4. For the path planning task between the warehouse and each car stop, a car path planning reinforcement learning model is constructed by designing the state space, action space and reward function;
[0012] S5. Based on the vehicle path planning reinforcement learning model, an Actor-Critic training method is adopted, and the UAV route planning strategy model is used to guide the optimization of the vehicle path, and the vehicle path planning strategy model is trained to form a vehicle-UAV joint delivery path planning model.
[0013] Furthermore, in S1, the process of car-drone scenario construction and task allocation is as follows:
[0014] S1.1. Generate a delivery environment dataset
[0015] Generate multiple distribution environment data sets based on parameters such as the number of customer nodes and the maximum load of drones. Each environment data in each distribution environment data set includes the locations of a warehouse node and several customer nodes and the demand attributes of the customer nodes, where the demand of the customer nodes is not greater than the maximum load of the drone.
[0016] S1.2. Initialize cluster centers
[0017] The number of clusters is determined to be k according to the number of car stops k; k coordinates are randomly selected as the initial cluster centers;
[0018] S1.3, calculate the Euclidean distance from each customer node to each cluster center in all distribution environment data sets, and assign each customer to the cluster center closest to it based on the calculated distance to form a preliminary clustering result;
[0019] S1.4. Update cluster centers
[0020] The mean of all customer node coordinates in each cluster is used as the new cluster center;
[0021] S1.5. Repeated iterative update
[0022] Repeat S1.3 to S1.4 until the cluster center no longer changes or the predetermined maximum number of iterations is reached;
[0023] S1.6. Get the delivery task
[0024] Each delivery environment outputs k clustering results. The center of each cluster is used as the location of the car stop, and the cluster of each cluster is used as the delivery area of the drone at the corresponding car stop, providing data support for subsequent car path planning and drone route planning.
[0025] Furthermore, in S2, the method for constructing a reinforcement learning model for drone route planning is as follows:
[0026] S2.1. Design state space
[0027] The state space is divided into a static part and a dynamic part; the static part includes the location of each node, including the assigned customer nodes and car stops; the dynamic part includes the location of the drone, the load, the remaining energy and the demand of each customer node; the load represents the amount of goods currently carried by the drone, the remaining energy represents the remaining power of the drone, and the demand of each customer node reflects how many items the customer node needs for delivery. The demand of the customer node becomes 0 after the delivery is completed, and the demand of the vehicle node reflects the negative value of the amount of goods delivered by the drone;
[0028] S2.2. Design Status Update
[0029] If the drone returns to the car stop, the load and remaining energy in the dynamic part of the state space are set to the maximum value; if the drone arrives at a customer node for delivery, the demand of the customer node is set to 0, indicating that the customer has completed the delivery, and the remaining load and remaining energy of the drone are updated based on the demand of the customer node and the route distance;
[0030] S2.3. Designing the Action Space
[0031] The action space includes selecting a new node from the current node for delivery, where the new node is a customer node or a car stop. In order to avoid invalid actions, a mask mechanism is introduced to prohibit the selection of customer nodes with a demand of 0 and customer nodes whose remaining load and energy of the drone cannot meet the demand or delivery distance.
[0032] S2.4. Designing reward functions
[0033] The reward function aims to minimize the energy consumption of the drone by penalizing the energy consumption of each delivery mission of the drone. The formula of the reward function is as follows:
[0034]
[0035] Among them, i, j∈N, N represents the set of car stops and customer nodes, c u represents the energy consumption of the drone per unit distance and per unit load, d ij Represents the distance between two nodes, l i represents the load of the drone at node i, y ij Indicates whether the route contains the path from node i to node j. If the drone route contains a path from node i to node j, then y ij =1, otherwise y ij =0;
[0036] Every time the drone flies from one node to another, it calculates the energy consumed based on the flight distance and load weight, ensuring that the drone chooses the most energy-efficient path when planning its route.
[0037] Furthermore, in S3, the training process of the drone route planning strategy model is as follows:
[0038] S3.1. Initialize reinforcement learning training parameters: Initialize the policy network (Actor) π UAV ,parameter And its learning rate
[0039] Initialize the value network (Critic) parameter And its learning rate
[0040] Initialize the reward accumulation value
[0041] Initialize the drone state: the node i where the drone is currently located is the car stop, and the remaining load is the maximum value l i =Q u , the remaining energy is the maximum value e i =E u ;
[0042] Initialize the path list and customer demand D = {d i}, i∈M, M represents the set of customer nodes, d i represents the demand of customer node i;
[0043] S3.2. Output action: observe the current state Make decisions based on the policy network (Actor) and generate actions That is, select the next step node j and let the drone perform the action, where
[0044] S3.3. Update state: observe new state from the environment Add node j to the path list L UAV ; If node j is a car stop, update the remaining load of the drone to the maximum value l j =Q u , the remaining energy is the maximum value e j =E u ; If node j is a client node, update the remaining load l of the drone j = l i -d j 、Residual energy e j =e i -e ij , the demand d of customer node j j Set to 0; e ij represents the energy consumed from node i to node j;
[0045] S3.4. Calculating Reward Values: Observing Instant Rewards from the Environment And update the total reward
[0046] S3.5. Make decisions based on the policy network (Actor): But don't let the drone perform actions
[0047] S3.6. Value Network (Critic) for Observed To rate:
[0048]
[0049] Actions that have not yet been executed as determined by the policy network And the observed state To rate:
[0050]
[0051] S3.7. Based on instant rewards and the value of the next state, calculate the temporal difference (TD) target and TD error
[0052]
[0053] Among them, γ UAV represents the discount rate for reinforcement learning;
[0054] S3.8. Update the value network (Critic): in Represents the gradient of the calculated value network parameters;
[0055] S3.9, Update strategy network (Actor): in Represents the gradient of the calculation policy network parameters;
[0056] S3.10, repeat S3.2-S3.9 until all customer requirements D are met and the UAV route planning strategy model is obtained;
[0057] S3.11. On the multiple delivery environment data sets in S1, the UAV route planning strategy model obtained in S3.10 is cyclically trained according to the training process from S3.1 to S3.10 to train and optimize the UAV route planning strategy model π UAV .
[0058] Furthermore, in S4, the method for constructing a vehicle path planning reinforcement learning model is as follows:
[0059] S4.1 Design state space: The state space includes the warehouse, the location of the car stop, and the identification of whether the car has visited;
[0060] S4.2 Design status update: After a car visits a car stop, the flag is set to 1 to avoid repeated selection;
[0061] S4.3 Design action space: The action space includes selecting the next node from the current bus stop, which can be a bus stop or a warehouse. In order to avoid invalid actions, a mask mechanism is introduced to prohibit the selection of bus stops that have been selected. When all bus stops have been selected, return to the warehouse to complete the delivery. Ensure that all service areas are visited and no loops are formed in the path.
[0062] S4.4 Design reward function: The reward function of the vehicle path planning mainly focuses on minimizing the total energy consumption of the vehicle and the drone. The reward function calculates the energy consumption based on the distance traveled from one stop to another and the transportation cost. At the same time, it calculates the energy consumption based on π UAV The generated drone route calculates the energy consumption of the drone UAV ;
[0063] r total =r Vechicle +r UAV
[0064]
[0065] Among them, r total represents the overall reward, r Vehicle represents the reward of the car, m, n∈K, K represents the set of car stops and warehouses, c v Indicates the energy consumption of the car per unit distance, d mn Represents the distance between two nodes, x mn Indicates whether the car path contains the path from node m to node n. If there is a path from node m to node n, then x mn =1, otherwise x mn =0.
[0066] Furthermore, in S5, the training process of the vehicle path planning strategy model is as follows:
[0067] S5.1. Initialize reinforcement learning training parameters: Initialize the policy network (Actor) π Vehicle ,parameter And its learning rate
[0068] Initialize the value network (Critic) parameter And its learning rate
[0069] Initialize the total reward accumulation value r total =0;
[0070] Initialize the car state: the node m where the car is currently located is the warehouse node;
[0071] Initialize the path list and the bus stop visit status P = {p m}, m∈K′, K′ represents the set of car stops, p m Indicates the visit status of the bus stop m;
[0072] S5.2. Output action: observe the current state Make decisions based on the policy network (Actor) and generate actions That is, select the next step node n and let the car perform the action, where And let the drone take off from the car stop, according to S3's drone route planning strategy model π UAV Generate delivery routes;
[0073] S5.3. Update state: observe new state from the environment Add node n to the path list L Vehicle ; Update access status pn =1;
[0074] S5.4. Calculating Reward Values: Observing Immediate Rewards from the Environment Calculating drone energy consumption And update the total reward
[0075] S5.5. Make decisions based on the policy network (Actor): But don't let the car perform actions
[0076] S5.6. Value Network (Critic) for Observed To rate: And the actions that have not yet been executed are decided by the policy network And the observed state To rate:
[0077]
[0078] S5.7. Calculate the temporal difference (TD) target based on the immediate reward and the value of the next state and TD error
[0079]
[0080] Among them, γ is the discount rate of reinforcement learning;
[0081] S5.8. Update the value network (Critic): in Represents the gradient of the calculated value network parameters;
[0082] S5.9. Update strategy network (Actor): in Represents the gradient of the calculation policy network parameters;
[0083] S5.10, repeat S5.2-S5.9 until all the car stops have been visited, and obtain the car path planning strategy model;
[0084] S5.11, on the multiple distribution environment data sets in S1, the vehicle path planning strategy model obtained in S5.10 is cyclically trained according to the training process from S5.1 to S5.10 to train and optimize the vehicle path planning strategy model π Vehicle , and the UAV route planning strategy model π UAVTogether they form a joint car-drone delivery path planning model.
[0085] Beneficial effects of the present invention:
[0086] 1) The present invention combines drone route planning with car path planning to fully utilize the collaborative advantages of cars and drones and improve overall delivery efficiency;
[0087] 2) The present invention uses reinforcement learning and Actor-Critic algorithm to optimize the path planning strategy, so that the distribution system can respond flexibly in a dynamic environment;
[0088] 3) The present invention uses the K-means clustering method to reasonably allocate customer nodes, ensuring the balance between the delivery efficiency and optimization effect of the vehicle-drone combined path planning model;
[0089] 4) The present invention is more in line with actual conditions by comprehensively considering the overall energy consumption of cars and drones and factors such as the load and energy limitation of drones, and can achieve efficient path planning in complex scenarios, thereby improving the overall performance of the logistics distribution system and reducing the overall energy consumption of cars and drones. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] Figure 1 It is the overall framework diagram of the logistics distribution path planning of the automobile-UAV of the present invention.
[0091] Figure 2 The present invention is a flow chart for calculating the allocation of car stops and drone customers based on the K-means clustering method.
[0092] Figure 3 This is a training structure diagram of the reinforcement learning model based on Actor-Critic of the present invention. DETAILED DESCRIPTION
[0093] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the specific implementation modes of the present invention will be described in detail below with reference to the accompanying drawings.
[0094] like Figure 1 As shown, an embodiment of the present invention provides a method for planning a vehicle-drone logistics distribution path based on deep reinforcement learning, comprising the following steps:
[0095] Step 1, such as Figure 1 As shown in Figure 2, a combined car-drone logistics delivery scenario is constructed, that is, multiple delivery environment datasets are generated. Figure 2As shown in the figure, tasks are assigned to cars and drones based on the locations of customer nodes, that is, the customer nodes in the distribution environment dataset are clustered using the K-means clustering method, each cluster center is used as the location of the car stop, and each cluster cluster is used as the delivery area of the drone at the corresponding car stop; the specific process is as follows:
[0096] 1.1) Generate distribution environment data set: It is set to generate 15 pieces of environmental data, each piece of environmental data contains 100 customer nodes, and the maximum load of the drone is 20. Therefore, each piece of environmental data generated contains the location information of 100 customer nodes and 1 warehouse node and the demand of customer nodes, where the demand of each customer node is a random number between 1 and 20.
[0097] 1.2) Initialize cluster centers: Set 10 car stops, so the corresponding number of clusters k = 10; randomly select 10 coordinates on the map as the initial cluster centers.
[0098] 1.3) Calculate the Euclidean distance from each customer node to all cluster centers to determine which cluster center the customer is closest to, and assign each customer to the nearest cluster center based on the calculated distance to form a preliminary clustering result.
[0099] 1.4) Update cluster centers: Take the mean of all client node coordinates in each cluster as the new cluster center.
[0100] 1.5) Repeat iterative update: Repeat steps 1.3 to 1.4 until the cluster center no longer changes or the predetermined maximum number of iterations is reached.
[0101] 1.6) Obtaining the delivery task: Each of the 15 delivery environments ultimately outputs 10 clustering results as the car’s stop points, corresponding to 10 delivery areas, providing data support for subsequent car path planning and drone route planning.
[0102] Step 2: Build a reinforcement learning model for the route planning task of the drone from the car stop to the customer node. The specific method is as follows:
[0103] 2.1) The state space of the drone model includes the locations of the customer nodes and car stops in the environment, the location of the drone, the load l, the remaining energy e, and the demand of each customer node D = {d i}, i∈M.
[0104] 2.2) The drone’s state update includes updating the remaining load and energy according to the selected action. If a car stop is selected, the remaining load is updated to the maximum value Q u =20, the remaining energy is the maximum value E u=100; if a customer node is selected, the remaining load and energy minus the corresponding consumption of the action, and the demand d of the selected customer node is updated j =0.
[0105] 2.3) The action space of the drone is to select the node i∈N to fly to next, where N represents the set of car stops and customer nodes.
[0106] 2.4) The reward function of the drone can be expressed as: r UAV =min E UAV =-∑ i∈N ∑ j∈N c u ·d ij ·l i ·y ij , where i, j∈N represents a car stop or customer node, c u represents the energy consumption of the drone per unit distance and per unit load, d ij Represents the distance between two nodes, l i represents the load of the drone at node i, y ij Indicates whether the route contains the path from node i to node j. If the drone route contains a path from node i to node j, then y ij =1, otherwise y ij =0.
[0107] Step three, such as Figure 3 As shown in the figure, the strategy network and value network are optimized by the Actor-Critic reinforcement learning method, and the UAV route planning strategy model 1π is trained and optimized. UAV ; The specific process is as follows:
[0108] 3.1) Initialize reinforcement learning training parameters: Initialize the policy network (Actor) π UAV ,parameter And its learning rate
[0109] Initialize the value network (Critic) parameter And its learning rate
[0110] Initialize the reward accumulation value
[0111] Initialize the drone state: the current node i is the car stop, and the remaining load is the maximum value Q u =20, the remaining energy is the maximum value E u =100;
[0112] Initialize the path list and customer demand D = {d i}, i∈M, M represents the set of customer nodes, d i represents the demand of customer node i;
[0113] 3.2) Output action: observe the current state Make decisions based on the policy network (Actor) and generate actions That is, select the next step node j and let the drone perform the action, where
[0114] 3.3) Update state: observe new state from the environment Add node j to the path list L UAV ; If node j is a car stop, update the remaining load of the drone to the maximum value l j =Q u , the remaining energy is the maximum value e j =E u ; If node j is a client node, update the remaining load l of the drone j = l i -d j 、Residual energy e j =e i -e ij , the demand d of customer node j j Set to 0; e ij represents the energy consumed from node i to node j;
[0115] 3.4) Calculate reward value: observe the immediate reward from the environment And update the total reward
[0116] 3.5) Make decisions based on the policy network (Actor): But don’t let the drone perform this action;
[0117] 3.6) The value network (Critic) has already observed a t 、s t To score, And the actions that have not yet been executed are decided by the policy network And the observed state To rate:
[0118]
[0119] 3.7) Calculate the temporal difference (TD) target based on the immediate reward and the value of the next state and TD error Among them, γ UAV represents the discount rate for reinforcement learning;
[0120] 3.8) Update the value network (Critic): in Represents the gradient of the calculated value network parameters;
[0121] 3.9) Update the policy network (Actor): in Represents the gradient of the calculation policy network parameters;
[0122] 3.10) Repeat steps 3.2) to 3.9) until all customer needs D are met, and the UAV route planning strategy model π is obtained. UAV1 ;
[0123] 3.11) Follow the training process from step 3.1) to step 3.10) to train the model π UAV1 The UAV route planning strategy model π is obtained by training on one of the 15 delivery environment datasets generated in step 1 (each dataset corresponds to a specific delivery scenario of 10 UAVs). UAV2 , and then for the model π UAV2 Another dataset from the 15 delivery environment datasets is used for training, and so on. After all 15 datasets are trained in a loop, the training optimizes the drone route planning strategy model π UAV .
[0124] Step 4: Build a vehicle path planning reinforcement learning model for the path planning task between the warehouse and each stop. The specific process is as follows:
[0125] 4.1) The state space of the car includes the location of the warehouse and the stop, and whether it has been visited before P = {p m}, m∈K′, K′ represents the set of car stops, p m Indicates the visit status of the bus stop m;
[0126] 4.2) The state of the car is updated as follows: after the car visits a car stop, the flag is set to p m =1, avoid repeated selection;
[0127] 4.3) The action space of the car is to select the next node, which is the car stop or warehouse. According to the mark, avoid choosing the stop that has been selected; when all the stops have been selected, that is, Then return to the warehouse to complete the delivery;
[0128] 4.4) The reward function of the car needs to minimize the total energy consumption of the whole, including the car and the drone. The reward function calculates the energy consumption based on the distance traveled from one stop to another and the transportation cost, and calculates the energy consumption through π UAV The generated drone route calculates the energy consumption of the drone UAV , so the reward function can be expressed as:
[0129] r total =r Vechicle +r UAV
[0130]
[0131] Among them, r total represents the overall reward, r Vehicle represents the reward of the car, m, n∈K, K represents the set of car stops and warehouses, c v Indicates the energy consumption of the car per unit distance, d mn Represents the distance between two nodes, x mn Indicates whether the car path contains the path from node m to node n. If there is a path from node m to node n, then x mn =1, otherwise x mn =0.
[0132] Step 5: Car-UAV combined delivery reinforcement learning model training. In the training of the car path planning strategy model, such as Figure 3 As shown in Figure 2, we construct an Actor-Critic based training method and use the drone strategy model π UAV Guide the optimization of the car path and train the car's path planning strategy model π Vehicle , forming the optimal path planning model for the joint vehicle-UAV; the specific method is as follows:
[0133] 5.1) Initialize reinforcement learning training parameters: Initialize the policy network (Actor) π Vehicle ,parameter And its learning rate
[0134] Initialize the value network (Critic) parameter And its learning rate
[0135] Initialize the total reward accumulation value r total =0;
[0136] Initialize the car state: the node m where the car is currently located is the warehouse node depot;
[0137] Initialize the path list and the bus stop visit status P = {p m}, m∈K′, K′ represents the set of car stops, p m Indicates the visit status of the bus stop m;
[0138] 5.2) Output action: observe the current state Make decisions based on the policy network (Actor) and generate actions That is, select the next step node n and let the car perform the action, where And let the drone take off from the stop point, according to the route planning strategy π in step three UAV Generate delivery routes;
[0139] 5.3) Update state: observe new state from the environment Add stop n to the path list L Vehicle ; Update access status p n =1;
[0140] 5.4) Calculating Reward Values: Observing Instant Rewards from the Environment Calculating drone energy consumption And update the total reward
[0141] 5.5) Make decisions based on the policy network (Actor): But don't let the car perform actions
[0142] 5.6) Value Network (Critic) for Observed To rate: And the actions that have not yet been executed are decided by the policy network And the observed state To rate:
[0143]
[0144] 5.7) Calculate the temporal difference (TD) target based on the immediate reward and the value of the next state and TD error Among them, γ is the discount rate of reinforcement learning;
[0145] 5.8) Update the value network (Critic): in Represents the gradient of the calculated value network parameters;
[0146] 5.9) Update the policy network (Actor): in Represents the gradient of the calculation policy network parameters;
[0147] 5.10) Repeat steps 5.2) to 5.9) until all car stops have been visited, and the car path planning strategy π is obtained. Vehicle1 ;
[0148] 5.11) Follow the training process from step 5.1) to step 5.10) to train the model π Vehicle1 The UAV route planning strategy model π is trained on one of the 15 delivery environment datasets generated in step 1 (each dataset corresponds to a specific delivery scenario of a car). Vehicle2 , in the model π Vehicle2 Another dataset from the 15 delivery environment datasets is used for training, and so on. After all 15 datasets are trained in a loop, the training optimizes the vehicle path planning strategy model π Vehicle , and the UAV route planning strategy π UAV Together they constitute a car-drone collaborative delivery method.
[0149] According to the above steps, the method of the present invention is compared with two classic heuristic algorithms, the ant colony algorithm and the simulated annealing algorithm, and the average planning time and average overall energy consumption of each algorithm are compared when there are 100 customer nodes and 200 customer nodes, respectively. As can be seen from Table 1, the overall energy consumption performance of the path plan planned by the method of the present invention is close to or slightly better than that of the classic algorithm; at the same time, since the method of the present invention adopts a reinforcement learning algorithm, the results can be quickly output through the trained strategy model, so it is significantly better than the heuristic algorithm in the calculation time of outputting the path planning results.
[0150] Table 1
[0151]
[0152] The above is a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A vehicle-drone logistics distribution path planning method based on deep reinforcement learning, characterized in that: The following steps are involved: S1. Generate a delivery environment data set, and cluster the customer nodes in the delivery environment data set using the K-means clustering method, with each cluster center serving as the location of a car stop, and each cluster cluster serving as the delivery area of the drone at the corresponding car stop; S2. For the UAV route planning task between the car stop and the customer node, a reinforcement learning model for UAV route planning is constructed by designing the state space, action space and reward function; S3. Based on the UAV route planning reinforcement learning model, the policy network and the value network are optimized by the Actor-Critic method to train and optimize the UAV route planning policy model; S4. For the path planning task between the warehouse and each car stop, a car path planning reinforcement learning model is constructed by designing the state space, action space and reward function; S5. Based on the vehicle path planning reinforcement learning model, an Actor-Critic training optimization method is adopted, and the UAV route planning strategy model is used to guide the optimization of the vehicle path, and the vehicle path planning strategy model is trained to form a vehicle-UAV joint delivery path planning model.
2. The method for automobile-UAV logistics distribution path planning based on deep reinforcement learning according to claim 1 is characterized in that: The specific process of S1 is as follows: S1.
1. Generate distribution environment data set: Generate multiple distribution environment data sets according to the number of customer nodes and the maximum load of the drone. Each environment data in each distribution environment data set includes the location of a warehouse node and several customer nodes and the demand attribute of the customer node, where the demand of the customer node is not greater than the maximum load of the drone. S1.2, Initialize cluster centers: Determine the number of clusters as k according to the number of car stops k; Randomly select k coordinates as initial cluster centers; S1.3, calculate the Euclidean distance from each customer node to each cluster center in all distribution environment data sets, and assign each customer to the cluster center closest to it based on the calculated distance to form a preliminary clustering result; S1.4, Update cluster center: take the mean of all customer node coordinates in each cluster as the new cluster center; S1.5, iterative update: repeat S1.3 to S1.4 until the cluster center no longer changes or the predetermined maximum number of iterations is reached; S1.
6. Get the delivery task: Output k clustering results for each delivery environment, each cluster center is used as the location of the car stop, and each cluster cluster is used as the delivery area of the drone at the corresponding car stop.
3. The method for automobile-UAV logistics distribution path planning based on deep reinforcement learning according to claim 1 or 2, characterized in that: In S2, the method for constructing a UAV route planning reinforcement learning model is as follows: S2.
1. Design state space The state space is divided into a static part and a dynamic part; the static part includes the location of each node, including the assigned customer nodes and car stops; the dynamic part includes the location of the drone, the load, the remaining energy and the demand of each customer node; the load represents the amount of goods currently carried by the drone, the remaining energy represents the remaining power of the drone, and the demand of each customer node reflects how many items the customer node needs for delivery, and the demand of the customer node becomes 0 after the delivery is completed; S2.
2. Design Status Update If the drone returns to the car stop, the load and remaining energy in the dynamic part of the state space are set to the maximum value; if the drone arrives at a customer node for delivery, the demand of the customer node is set to 0, indicating that the customer has completed the delivery, and the remaining load and remaining energy of the drone are updated based on the demand of the customer node and the route distance; S2.
3. Designing the Action Space The action space includes selecting a new node from the current node for delivery, where the new node is a customer node or a car stop. In order to avoid invalid actions, a mask mechanism is introduced to prohibit the selection of customer nodes with a demand of 0 and customer nodes whose remaining load and energy of the drone cannot meet the demand or delivery distance. S2.
4. Designing reward functions The reward function aims to minimize the energy consumption of the drone by penalizing the energy consumption of each delivery mission of the drone. The formula of the reward function is as follows: Among them, i,j∈N, N represents the set of car stops and customer nodes, c u represents the energy consumption of the drone per unit distance and per unit load, d ij Represents the distance between two nodes, l i represents the load of the drone at node i, y ij Indicates whether the route contains the path from node i to node j. If the drone route contains a path from node i to node j, then y ij =1, otherwise y ij =0.
4. The method for automobile-UAV logistics distribution path planning based on deep reinforcement learning according to claim 3 is characterized in that: In S3, the training process of the drone route planning strategy model is as follows: S3.
1. Initialize reinforcement learning training parameters: Initialize the policy network π UAV ,parameter And its learning rate Initialize the value network parameter And its learning rate Initialize the reward accumulation value Initialize the drone state: the node i where the drone is currently located is the car stop, and the remaining load is the maximum value l i =Q u , the remaining energy is the maximum value e i =E u ; Initialize the path list and customer demand D = {d i }, i∈M, M represents the set of customer nodes, d i represents the demand of customer node i; S3.
2. Output action: observe the current state Make decisions based on the policy network and generate actions That is, select the next step node j and let the drone perform the action, where S3.
3. Update state: observe new state from the environment Add node j to the path list L UAV ; If node j is a car stop, update the remaining load of the drone to the maximum value l j =Q u , the remaining energy is the maximum value e j =E u ; If node j is a client node, update the remaining load l of the drone j = l i -d j 、Residual energy e j =e i -e ij , the demand d of customer node j j Set to 0; e ij represents the energy consumed from node i to node j; S3.
4. Calculating Reward Values: Observing Instant Rewards from the Environment And update the total reward S3.
5. Make decisions based on the policy network: But don't let the drone perform actions S3.
6. Value Network for Observed To rate: Actions that have not yet been executed as determined by the policy network And the observed state To rate: S3.
7. Based on instant rewards and the value of the next state, calculate the time difference target and time difference error Among them, γ UAV represents the discount rate for reinforcement learning; S3.
8. Update value network: in Represents the gradient of the calculated value network parameters; S3.9, Update strategy network: in Represents the gradient of the calculation policy network parameters; S3.10, repeat S3.2-S3.9 until all customer requirements D are met and the UAV route planning strategy model is obtained; S3.
11. On multiple distribution environment data sets, the UAV route planning strategy model obtained in S3.10 is cyclically trained according to the training process from S3.1 to S3.10 to train and optimize the UAV route planning strategy model π UAV .
5. The method for automobile-UAV logistics distribution path planning based on deep reinforcement learning according to claim 4 is characterized in that: In S4, the method for constructing the vehicle path planning reinforcement learning model is as follows: S4.1 Design state space: The state space includes the warehouse, the location of the car stop, and the identification of whether the car has visited; S4.2 Design status update: After a car visits a car stop, the flag is set to 1 to avoid repeated selection; S4.3 Design action space: The action space includes selecting the next node from the current bus stop, which can be a bus stop or a warehouse. In order to avoid invalid actions, a mask mechanism is introduced to prohibit the selection of bus stops that have been selected. When all bus stops have been selected, return to the warehouse to complete the delivery. Ensure that all service areas are visited and no loops are formed in the path. S4.4 Design reward function: The reward function of the vehicle path planning focuses on minimizing the total energy consumption of the vehicle and the drone. The reward function calculates the energy consumption based on the distance traveled from one stop to another and the transportation cost. UAV The generated drone route calculates the energy consumption of the drone UAV ; r total =r Vechicle +r UAV Among them, r total represents the overall reward, r Vehicle represents the reward of the car, m,n∈K, K represents the set of car stops and warehouses, c v Indicates the energy consumption of the car per unit distance, d mn Represents the distance between two nodes, x mn Indicates whether the car path contains the path from node m to node n. If there is a path from node m to node n, then x mn =1, otherwise x mn =0.
6. The method for automobile-UAV logistics distribution path planning based on deep reinforcement learning according to claim 5 is characterized in that: In S5, the training process of the vehicle path planning strategy model is as follows: S5.
1. Initialize reinforcement learning training parameters: Initialize the policy network π Vehicle ,parameter And its learning rate Initialize the value network parameter And its learning rate Initialize the total reward accumulation value r total =0; Initialize the car state: the node m where the car is currently located is the warehouse node; Initialize the path list and the bus stop visit status P = {p m }, m∈K', K' represents the set of car stops, p m Indicates the access status of the bus stop m; S5.
2. Output action: observe the current state Make decisions based on the policy network and generate actions That is, select the next step node n and let the car perform the action, where And let the drone take off from the car stop, according to S3's drone route planning strategy model π UAV Generate delivery routes; S5.
3. Update state: observe new state from the environment Add node n to the path list L Vehicle ; Update access status p n =1; S5.
4. Calculating Reward Values: Observing Immediate Rewards from the Environment Calculating drone energy consumption And update the total reward S5.
5. Make decisions based on the policy network: But don't let the car perform actions S5.
6. Value Network for Observed To rate: And the actions that have not yet been executed are decided by the policy network And the observed state To rate: S5.
7. Calculate the time difference target based on the immediate reward and the value of the next state and time difference error Among them, γ is the discount rate of reinforcement learning; S5.
8. Update value network: in Represents the gradient of the calculated value network parameters; S5.
9. Update strategy network: in Represents the gradient of the calculation policy network parameters; S5.10, repeat S5.2-S5.9 until all the car stops have been visited, and obtain the car path planning strategy model; S5.
11. On multiple distribution environment data sets, the vehicle path planning strategy model obtained in S5.10 is cyclically trained according to the training process from S5.1 to S5.10 to train and optimize the vehicle path planning strategy model π Vehicl , and the UAV route planning strategy model π UAV Together they form a joint car-UAV delivery path planning model.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor; characterized in that: When the processor executes the computer program, the electronic device executes the car-drone logistics distribution path planning method based on deep reinforcement learning as described in any one of claims 1-6.
8. A storage medium, comprising a computer program, which, when executed on an electronic device, enables the electronic device to execute the deep reinforcement learning-based automobile-drone logistics distribution path planning method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Path planning method and system based on reinforcement learning and heuristic search
CN111896006A
Path planning method based on improved deep reinforcement learning
CN112362066A
Mobile robot path planning method based on improved DDPG algorithm
CN114089751A
Water-air amphibious unmanned vehicle path planning method based on reinforcement learning
CN114089762A
Heterogeneous vehicle type vehicle path planning method and system based on deep reinforcement learning
CN118608021A