An urban logistics distribution method based on the collaboration of public transportation and drones

The two-layer decision-making framework constructed through the Dueling DQN algorithm with a multi-head attention mechanism and a genetic algorithm solves the problem of drone path planning relying on bus trajectories and insufficient package matching, realizes an efficient and green logistics model of collaborative delivery between buses and drones, and improves the efficiency of urban logistics distribution.

CN120579915BActive Publication Date: 2025-09-30DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511083206.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-09-30
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

In existing technologies, drone path planning is overly dependent on bus driving trajectories and cannot be flexibly adjusted. There is also a lack of systematic solutions to match packages with bus schedules, resulting in limited collaborative delivery efficiency.

Method used

The Dueling DQN algorithm integrated with the multi-head attention mechanism is used to make matching decisions between packages and bus schedules. The genetic algorithm is combined with the UAV path planning to construct a two-layer decision framework to achieve reasonable matching between packages and buses and optimization of UAV paths.

Benefits of technology

It has improved delivery efficiency, revitalized public transportation resources, reduced traffic congestion and pollution, realized an efficient network for coordinated delivery between public transportation and drones, and enhanced the green sustainability of urban logistics delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579915B_ABST
    Figure CN120579915B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of urban smart logistics technology and relates to an urban logistics distribution method that collaborates with buses and drones. The present invention is a new bus-drone collaborative distribution model suitable for non-peak traffic scenarios, and designs and constructs a two-layer decision-making framework. The upper-layer decision-making adopts the Dueling DQN algorithm that integrates a multi-head attention mechanism to reasonably match the package with the bus it should take. After the package is transported to a specific drone station by bus according to the upper-layer decision-making result, the lower-layer decision-making uses a heuristic algorithm to plan the paths of each drone. The present invention forms an efficient network of "trunk transportation + flexible terminals" by coordinating the idle transportation capacity of buses during non-peak periods with drone terminal distribution, which not only activates bus resources but also improves distribution efficiency, reducing traffic and pollution problems caused by additional vehicles in traditional solutions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of urban smart logistics technology and relates to an urban logistics distribution method that collaborates with public transportation and drones. Specifically, it comprises an urban logistics system architecture that integrates idle public transportation capacity with collaborative drone distribution, as well as a two-layer decision-making optimization method for matching packages with public transportation schedules and planning drone paths. The method aims to improve urban logistics distribution efficiency and promote the development of green logistics through cross-domain technology integration. Background Art

[0002] The recent boom in e-commerce has posed significant distribution challenges for logistics service providers. Under the traditional "truck + manual" model, urban truck traffic has surged, exacerbating congestion and contributing to traffic chaos, pollution, and noise. Furthermore, traditional ground buses in many cities are facing a severe loss of passenger traffic, resulting in significant waste of off-peak bus capacity and heavy losses for bus companies. Traditional buses are facing a severe survival crisis. The development of a new, green, efficient, and sustainable urban logistics model is imperative.

[0003] With the continuous advancement of drone technology, drone-assisted delivery has become a research hotspot in the logistics field due to its environmental, efficiency, and energy-saving advantages. However, due to limitations in payload capacity and battery capacity, drones cannot independently complete long-distance, heavy-load delivery missions. Although existing research has attempted to address this issue through the coordinated operation of trucks and drones, the additional truck resources still cause secondary problems such as traffic congestion and environmental pollution. In contrast, urban public transportation networks have the advantage of independent road rights and existing public transportation infrastructure, eliminating the need for reconstruction and causing additional traffic congestion and pollution.

[0004] In research on integrating drone delivery into public transportation systems, Peng Yong et al. (CN115423406A) disclosed a drone-bus regional collaborative delivery system and method. The drones are responsible for delivery and storage, the buses transport and charge the drones, and a smart terminal app handles order placement and display. However, the drones in this method are overly dependent on bus routes, requiring them to take specific routes or take off and land at specific locations, making it difficult to flexibly adapt delivery to actual conditions. Zhang Peng et al. (CN117094621A) disclosed an urban logistics transportation method based on drone-bus collaboration. The drones transport packages to the top of the bus, which then travels with the packages, and the drones pick up and deliver the packages. This method suffers from the following drawbacks: it only considers the process design for drone-bus collaboration, but does not address bus capacity allocation, nor does it consider how each package selects its specific bus schedule (e.g., how packages are allocated based on bus load, route, arrival time, and other factors). Zhang Qi et al. disclosed a bus network-assisted drone delivery scheduling method, system, and storage medium (CN118229181B). Drones can choose buses on connecting bus routes for onboard transportation. However, in the intermodal bus mode, the drone path is limited by bus stops and trajectory data, and the drone cannot perform flexible delivery according to actual conditions.

[0005] In the research on bus and drone collaborative delivery, existing technologies still have significant shortcomings:

[0006] 1) Drone path planning is overly dependent on bus trajectories, allowing drones to take off and land only on specific routes or at specific stops, making it difficult to flexibly adjust the optimal path based on actual delivery needs.

[0007] 2) Existing methods lack a systematic solution for matching packages with bus schedules and do not involve capacity allocation strategies based on multi-dimensional factors such as bus load, route planning, and arrival time, resulting in limited collaborative delivery efficiency. Summary of the Invention

[0008] Based on the above problems, the present invention proposes a new bus-UAV collaborative delivery model suitable for non-peak traffic scenarios, and designs and constructs a two-layer decision-making framework. The upper-layer decision-making adopts the Dueling DQN algorithm integrated with the multi-head attention mechanism to reasonably match the package with the bus it should take. After the package is transported to a specific UAV station by bus according to the upper-layer decision-making result, the lower-layer decision-making uses a genetic algorithm to plan the path of each UAV.

[0009] The technical solutions of the present invention are as follows:

[0010] A method for urban logistics distribution using public transportation and drones, including the following steps:

[0011] Step 1: Construction of urban logistics distribution system and basic data collection. The specific process is as follows:

[0012] Warehouses will be built at bus terminals on the city's outskirts, integrating warehousing and bus dispatch functions for package collection, distribution, and loading. Bus stops along bus routes located in areas with historically high parcel delivery demand will be upgraded to drone stations. Each station will be equipped with multiple drones (necessary to meet peak delivery demand) and battery swapping facilities to support rapid battery replacement upon return.

[0013] A cargo platform is installed on the roof of each bus. When the bus stops at its origin or destination (warehouse), packages are loaded onto the roof before departure. Buses depart according to a fixed schedule and travel along fixed routes. When a bus reaches a regular bus stop (not a drone stop), only passengers board and disembark. When a bus reaches a drone stop, not only do passengers board and disembark, but drones also remove packages from the bus that have reached their destination for delivery. When the bus stops at the terminal, packages are reloaded and the bus resumes its journey according to the schedule.

[0014] Obtain information about each bus schedule, drones, and packages that need to be delivered over the next period. Based on the minimum distance between the package's destination and the drone station, select a drone station as the destination for each package. Number each package and select the bus schedule it should take, starting with the smallest number and ending with the largest. Then, group the packages into groups based on the drone station they are destined to and the time it will arrive at the station.

[0015] Step 2: Model the bus and drone collaborative delivery problem as a Markov decision process. Build an upper-level package-bus schedule matching decision module. Based on the Dueling DQN algorithm with a multi-head attention mechanism, match packages with specific bus schedules. This means selecting the bus that each package should take. The specific process is as follows:

[0016] Step 2.1: State and action definition

[0017] The status information designed by the present invention is the information of each bus and the information of the package being distributed when distributing the current package. The status information is represented as ,in, Bus feature information, indicating the distribution of packages All The existing load situation of buses on the schedule, Package feature information, which represents the five-dimensional information of the package being distributed: number, weight, delivery time window (start time and end time), and rewards and penalties.

[0018] The action of the present invention Corresponding to the choice of allocating packages to different buses, specifically allocating the current package to One of the bus schedules, Indicates that the package Assigned to bus routes Take a ride.

[0019] Step 2.2: Multi-head attention mechanism processes state information

[0020] because and The correlation between these two features is not obvious, and the traditional Q network is difficult to effectively capture key information, so it is necessary to +5-dimensional state information is decoupled, projected, reorganized, and processed with attention.

[0021] First, decouple the input state information into dimensional bus feature information and the last 5-dimensional package feature information. Two projection layers are used to project the bus and package features into a 128-dimensional embedding space, respectively, so that the features have richer expressive power. Next, the projected bus and package features are concatenated in the sequence dimension to form a sequence with a shape of [2, batch, 128]. Batch is the batch size, that is, the number of experience samples sampled from the memory pool each time. This sequence is input into the multi-head attention layer to capture the association between state information from different subspaces in parallel. In the attention calculation process, this sequence is also used as the query ( )、key( ) and value ( ), calculates the attention weight for the interaction between bus and package features. The calculation formula is as follows:

[0022]

[0023] in, express The dimension size of .

[0024] After processing by the multi-head attention layer, the output attention result is shaped like [2, batch, 128]. The attention results of the bus feature information and the package feature information are then concatenated and converted to a shape of [batch, 256]. The concatenated features are input into the subsequent shared layer, where fused features are further extracted through a multi-layer fully connected network.

[0025] Step 2.3: Reward Function Design

[0026] The present invention realizes the combination of immediate and delayed rewards by constructing a linkage mechanism between a temporary memory bank and a memory pool.

[0027] When the upper layer decides to assign a bus to a package, the immediate reward is calculated on the fly by seeing if the package is assigned to a bus that is overloaded or will not reach the corresponding drone station. , and include the current state ,action , instant rewards and the next state The quadruple Stored in temporary memory, because the lower-level drone path planning is not yet completed, delaying the reward It cannot be calculated, so the temporary memory bank plays the role of temporarily storing real-time information in the decision-making process to ensure that the intermediate state of each allocation action can be recorded.

[0028] When the lower-level decision-making completes the drone path planning through the genetic algorithm, the rewards and penalties for whether each package is delivered on time, the cost of the drone flight distance, and the cost of the drone sorties are added together as the operating income of the group, and the operating income of all groups is calculated. and individual package revenue rewards , divide the two equally and take the weighted sum as the delayed reward Extract the corresponding package quad from the temporary memory bank and immediately reward With delayed rewards The sum of the two is the final reward , and update the quadruple Stored in the memory pool for subsequent model training. The calculation formula is as follows:

[0029]

[0030]

[0031] in, It is the weight parameter used to adjust the individual revenue reward of each package in the delayed reward. Its function is to balance the impact of individual package delivery results on the overall delayed reward. The quantity of all packages.

[0032] Step 2.4, Prioritize Experience Replay

[0033] In order to improve the learning efficiency of the intelligent agent for key state information and alleviate the problem of dilution of important experience caused by traditional uniform sampling, a priority experience replay mechanism is introduced. Through the dual mechanisms of non-uniform sampling and deviation correction, focused learning of high-value experience is achieved.

[0034] When new experience When stored in the memory pool, its sampling priority is initialized to 1.0 to ensure uniform exploration in the initial stage. As training progresses, the time difference is calculated at each experience replay ( )error:

[0035]

[0036] in, are the parameters of the current Q network (online network), are the parameters of the target Q network, Indicates the next state The possible actions to be taken. is the current network output, Output of the target network. The error reflects the contribution of experience to the update of the value function. The larger the error, the more important the experience. The error plus a small constant of 10 -5 To avoid zero priority, update the priority of the corresponding experience:

[0037]

[0038] in, To indicate the The probability that an experience is sampled, Indicates the The time difference error of the experience.

[0039] By exponential transformation ( is a hyperparameter used to adjust the experience priority weight) amplifies the priority difference and forms a non-uniform sampling probability distribution, which can make the sampling probability of high-error experience significantly higher than that of low-error experience:

[0040]

[0041] in, To indicate the The probability that an experience is sampled, It is the index of the experience sample, which traverses all the experience samples in the memory pool and is used to calculate the sampling probability When the priority of all experiences is summed, To indicate the Priority of experience of Power ( is a hyperparameter used to amplify priority differences).

[0042] At the same time, in order to offset the distribution bias introduced by priority sampling, the importance sampling weight is introduced :

[0043]

[0044] in, is the memory pool capacity, is a weight parameter used to correct the bias caused by non-uniform sampling.

[0045] Finally, according to the importance sampling weight Calculate the batch average loss by square the time difference error of each sample , and then the gradient is calculated by the back-propagation algorithm, and the optimizer updates the neural network parameters:

[0046]

[0047] in, is the batch size, that is, the number of experience samples sampled from the experience replay buffer each time.

[0048] The combination of non-uniform sampling and bias correction mechanisms enables the parameter update direction to benefit from important experience while not deviating from the true expectation of the overall state distribution, which can effectively solve the problem of low learning efficiency of traditional uniform sampling.

[0049] Step 3: Build a lower-level decision-making module to solve the drone delivery scheduling problem of the parcel group based on the genetic algorithm, and implement drone sortie planning, parcel loading combination optimization, and delivery route design for each group. The specific process is as follows:

[0050] First, the candidate routes are initialized. The destination stations are randomly arranged as the initial solution sequence, presented as an integer sequence of length X+2, consisting of two 0s and X package destinations. 0 represents the departure and return drone stations, and integers from 1 to X correspond to the delivery destinations, where X is the number of packages that the drone station must deliver. The last digit is also 0, indicating that the drone must depart from and return to the drone station. For example, the initial solution sequence [0, 6, 3, 1…z, 0], where the first digit is 0, 6, 3, 1…z, totaling X digits, and the last digit is also 0, indicates that the route departs from the drone station, delivers packages 6, 3, 1…z, and finally returns to the drone station.

[0051] By traversing the intermediate package destination nodes in the complete path represented by the solution sequence, for each node, the corresponding package weight is obtained, and a temporary path is constructed that includes the previous path, the node, and the return station, and the distance between them is calculated. Then, the drone is checked to see if the drone is overloaded after adding this package, and whether the temporary path distance exceeds the maximum flight distance. If any of these conditions are met, the current path plus the return station is stored as a complete path in the result list, and a new path is started. If neither condition is met, the node is added to the current path and the payload is updated. After the traversal is completed, the final path is added to the return station and stored in the result list, and the segmented drone path information is finally returned.

[0052] The fitness of each individual is calculated based on the total flight distance after path segmentation and the number of drones used. The individual's strengths are assessed and the best individuals are selected for the next generation. A crossover operation is performed to select parents from the best individuals, and genes are randomly exchanged at crossover points to generate offspring. Mutation is performed on individuals with a probability μ, modifying certain genes in the newborn individuals to increase population diversity. After the final iteration, the individuals with the best fitness are selected for path segmentation decoding, resulting in a final path plan for each package group that meets the drone payload and flight distance constraints.

[0053] Step 4: Model training

[0054] Initialize the Dueling DQN network with a multi-head attention mechanism, set the input state dimension and output action dimension, use the Adam optimizer and Step LR scheduler, and initialize the memory pool capacity. Allocate buses in order of package numbers using the ε-greedy strategy, calculate the immediate reward and store it in a temporary memory bank, and then merge it with the delayed reward and store it in the memory pool. Sample data from the memory pool according to the TD error priority, and update the network by calculating the weighted MSE loss using the importance weights. Regularly synchronize the target network parameters, and save the model after reaching the maximum number of iterations.

[0055] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0056] 1) By synergizing idle bus capacity during off-peak hours with drone terminal delivery, an efficient "trunk transport + flexible terminal" network is formed, which not only activates bus resources but also improves delivery efficiency, reducing traffic and pollution problems caused by additional vehicles in traditional solutions.

[0057] 2) A two-stage decision-making approach is constructed for the bus-drone collaborative delivery system proposed in this invention. The upper-level decision-making considers the lower-level path solution when matching packages with bus rides. The lower-level decision-making solves the delivery path for each package group, achieving coordinated optimization of delivery revenue and drone maintenance costs.

[0058] 3) Design an immediate-delayed reward linkage mechanism. Immediate rewards constrain the rationality of allocation, and delayed rewards are dynamically updated based on the final benefits of lower-level decisions. By linking the temporary memory bank with the memory pool, the decision chain is extended to the entire "allocation-execution" cycle, solving the short-sightedness problem of traditional reinforcement learning.

[0059] 4) Design a solution algorithm that integrates the multi-head attention mechanism and Dueling DQN. By decoupling, projecting, and reorganizing the heterogeneous information of bus load and package characteristics, it achieves adaptive weighted fusion of high-dimensional state information, significantly improving the ability to capture complex interactions and overcoming the inefficiency of traditional algorithms in processing high-dimensional state information. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 Schematic diagram of the logistics distribution system of the present invention;

[0061] Figure 2 It is a solution flow chart of the present invention;

[0062] Figure 3 It is a schematic diagram of the framework of the present invention;

[0063] Figure 4 Solve the delivery path diagram for each package group for the lower-level decision genetic algorithm;

[0064] Figure 5 The total revenue obtained by this method, DQN, and Dueling DQN algorithms when the demand for packages is 100.

[0065] Figure 6 The delivery rate obtained by this method, DQN, and Dueling DQN algorithms when the package demand is 100.

[0066] Figure 7 The total revenue obtained by this method, DQN, and Dueling DQN algorithms when the package demand is 150.

[0067] Figure 8 This is the delivery rate obtained by this method, DQN, and Dueling DQN algorithms when the package demand is 150. DETAILED DESCRIPTION

[0068] The objectives and technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0069] The present invention is a method for urban logistics distribution that is coordinated by public transportation and drones. Figure 2 The specific steps are as follows:

[0070] Step 1: Construction of urban logistics distribution system and basic data collection. The specific process is as follows:

[0071] Step 1.1, Infrastructure deployment, see Figure 1 :

[0072] Warehouses are built near several bus terminals around the city, integrating the terminals with the warehouses. All packages with addresses in the city are first transported to the designated warehouses for distribution. Each bus departs from the starting station (warehouse) with several packages on board for delivery.

[0073] Based on historical package delivery demand, drone stations have been constructed at several bus stops along various bus routes, integrating bus stops with drone stations. These stations are dedicated hubs for buses to pick up and drop off passengers and for drones to launch and land for pickup and delivery. Each station is equipped with several drones (necessary to meet peak delivery demand) and battery swap facilities to enable rapid recharging upon return. When a bus stops at a drone station, passengers board and exit, and a drone in the station takes off, selects a specific package on the bus, delivers it along a pre-planned route, and returns to the station.

[0074] A cargo platform is installed on the roof of each bus. When the bus stops at its origin or destination (warehouse), several packages are loaded onto the rooftop platform before departure. The bus departs according to a fixed schedule and travels along a fixed route. When the bus reaches a regular bus stop, only passengers board and disembark. When the bus reaches a drone stop, not only do passengers board and disembark, but drones also remove specific packages from the bus for delivery. When the bus stops at the terminal, packages are reloaded and the bus resumes its journey according to the schedule.

[0075] Step 1.2, basic data collection:

[0076] Determine the following system parameters:

[0077] Bus parameters: bus numbers for each bus , maximum load capacity , various bus routes Arrival at each drone station Time ;

[0078] Drone parameters: Maximum payload , Maximum flight distance , average flight speed ;

[0079] Package parameters: the number of each package ,weight , destination coordinates , delivery time window , delivery reward Punishment for failure to meet deadlines .

[0080] For each package , calculate the Euclidean distance between it and all drone stations and assign it to the drone station closest to it , forming an ownership relationship. Each package must take the bus to the drone station to which it belongs, and then be picked up by the drone for delivery. Arriving at the same drone station The packages are divided into a package group .

[0081] Step 2: Model the bus and drone collaborative delivery problem as a Markov decision process, build an upper-level package-bus schedule matching decision module, and use the Dueling DQN algorithm with a multi-head attention mechanism to achieve a reasonable match between packages and specific bus schedules. The specific process is as follows:

[0082] Step 2.1: Define states and actions

[0083] The status information designed by the present invention is the information of each bus and the information of the package being distributed, that is, In this embodiment, the bus operation frequency for the next period of time is set to 16. Indicates the distribution of packages The load conditions of all 16 buses at that time, , Indicates the number, weight, delivery time window (start time and end time) and reward and penalty information of the package being distributed. The order of allocating packages is based on the package numbers.

[0084] The action space designed by the present invention corresponds to the choice of allocating packages to different buses, specifically allocating the current package to be allocated to one of the 16 buses, which can be expressed as , which means to package Assigned to buses Carry on board.

[0085] Step 2.2: The multi-head attention mechanism processes state information, such as Figure 3 As shown in Section S2.2:

[0086] because and The correlation between these two parts of features is not obvious, and the traditional Q network is difficult to effectively capture key information. Therefore, it is necessary to decouple, project, reorganize and process the 21-dimensional state information.

[0087] First, the input state information is split into the current load information of each bus in the first 16 dimensions and the package features being distributed in the last 5 dimensions. In order to enable the model to better handle these two parts of features, this embodiment sets the use of two projection layers to project the features of the bus and the package into a 128-dimensional embedding space respectively, so that the features have richer expressive capabilities. Next, the projected bus and package features are spliced ​​in the sequence dimension to form a sequence with a shape of [2, batch, 128]. This sequence is input into the multi-head attention layer. In this embodiment, the number of heads M=8 is set to capture the association between the state information in parallel from 8 different subspaces. During the attention calculation process, the sequence is used as query, key and value at the same time, and the model will calculate the attention weight for the interaction between the bus and package features. The calculation formula is as follows:

[0088]

[0089] After processing by the multi-head attention layer, the output attention result is shaped like [2, batch, 128]. The attention results of the bus feature information and the package feature information are then concatenated and converted to a shape of [batch, 256]. This concatenated feature is then input into the subsequent shared layer, where it is further extracted through a multi-layer fully connected network to extract fused features.

[0090] Step 2.3, reward function design, such as Figure 3 As shown in Section S2.3:

[0091] This invention combines immediate and delayed rewards by building a linkage mechanism between a temporary memory bank and a memory pool. During the package distribution process, immediate rewards provide feedback based on the rationality of the current package distribution, constraining distribution behavior in real time. Delayed rewards, on the other hand, rely on the final outcome of lower-level decisions, taking into account factors such as whether each group's packages are delivered on time and the cost of drone flights to update the reward value.

[0092] The immediate reward includes penalties for overloading and missing buses. The delayed reward is based on the final group benefit and individual benefit feedback, balancing short-term and long-term decisions. If the current package is assigned to a bus and the bus is overloaded, or the package is assigned to a bus that will not reach its drone station, a penalty will be given, otherwise a reward will be given. Calculate the immediate reward After that, it will contain the current state , the actions taken , instant rewards and the next state The quadruple Store in temporary memory.

[0093] After the lower-level decision module uses the genetic algorithm to output the results, the rewards and penalties for whether each package is delivered on time, the cost of drone flight distance and the cost of drone sorties are added together as the operating income of the group. The total operating income of all groups is calculated. and the individual revenue value of each package After equal distribution and weighted adjustment, it is used as a delayed reward : , this embodiment sets the weight parameter , The total number of packages that need to be delivered.

[0094] Instant rewards will be given in the end With delayed rewards Add up, that is, take action in state s The reward r, In the reward After the update, the experience Stored in the memory pool.

[0095] Step 2.4, Prioritize Experience Replay:

[0096] In order to improve the learning efficiency of the intelligent agent for key states and alleviate the problem of dilution of important experience caused by traditional uniform sampling, the present invention introduces a priority experience replay mechanism, which realizes the focused learning of high-value experience through the dual mechanisms of non-uniform sampling and deviation correction.

[0097] When new experience When stored in the memory pool, its sampling priority is initialized to 1.0 to ensure uniform exploration in the initial stage. As training progresses, the temporal difference (TD) error is calculated at each experience replay:

[0098]

[0099] in, is the current network output, is the target network output. TD error reflects the contribution of experience to the update of the value function. The larger the error, the more important the experience. Add a small constant 10 to the TD error. -5 To avoid zero priority, update the priority of the corresponding experience:

[0100]

[0101] By exponential transformation ( It is a hyperparameter used to adjust the experience priority weight. ) Amplify the priority difference and form a non-uniform sampling probability distribution, which can make the sampling probability of high-error experience significantly higher than that of low-error experience:

[0102]

[0103] At the same time, in order to offset the distribution bias introduced by priority sampling, the present invention introduces importance sampling weights :

[0104]

[0105] in, is the memory pool capacity, is a weight parameter used to correct the deviation caused by non-uniform sampling (this embodiment adopts the annealing strategy, the initial , gradually rising to ).

[0106] Finally, according to the importance sampling weight Calculate the batch average loss by square the time difference error of each sample , and then the gradient is calculated by the back-propagation algorithm, and the optimizer updates the neural network parameters:

[0107]

[0108] in, is the batch size, that is, the number of experience samples sampled from the experience replay buffer each time.

[0109] The combination of non-uniform sampling and bias correction mechanisms enables the parameter update direction to benefit from important experience while not deviating from the true expectation of the overall state distribution, which can effectively solve the problem of low learning efficiency of traditional uniform sampling.

[0110] According to the decision of the upper layer, each package chooses the time to arrive at the drone station after taking the bus, and the same time t arrives at the same drone station The packages are divided into a package group After the upper level decision makers have distributed all the reports, the final package information of all package groups can be obtained.

[0111] Step 3: Construct the lower-level decision-making module and solve the drone delivery scheduling problem of the package group based on the genetic algorithm to realize the planning of drone sorties, package loading combination optimization and delivery path design for each group, such as Figure 4 As shown. The delay reward is calculated based on the final drone path results of each package group. , according to the reward function design in step 2.3, combined with the instant reward obtained previously, the reward r for taking action a in state s is calculated and stored in the memory pool for model training. The specific process is:

[0112] Step 3.1: Individual encoding and initialization

[0113] Package Group The delivery path is encoded as a sequence of integers, with 0 representing the drone station responsible for delivering the group's packages and 1 through X representing the package destinations. For example, [0, 6, 3, 1, ..., X, 0] means starting from the drone station, delivering packages 6, 3, 1, ..., X, and finally returning to the drone station. The initial population is generated by randomly permuting the destination points. In this example, 20 individuals are generated.

[0114] Step 3.2: Path segmentation and constraint processing

[0115] Since a single drone has load and flight distance limitations and cannot complete all package deliveries at once, the present invention uses a path segmentation method to solve the final path. By traversing the intermediate package nodes in the complete path represented by the initial solution sequence, for each package node, the corresponding package weight is obtained, a temporary path containing the node and the drone station is constructed and its distance is calculated. Then, it is checked whether the drone is overloaded after adding this package, and whether the temporary path distance exceeds the maximum flight distance. If any of the above situations exist, the current path plus the returned drone station is stored as a complete path in the result list, and a new path is opened; if neither exists, the node is added to the current path and the drone load is updated. After the traversal is completed, the return station is added to the last path and stored in the result list, and finally the segmented drone path information list is returned.

[0116] Step 3.3, genetic algorithm iteration

[0117] The fitness of each individual is calculated based on the total flight distance after path segmentation and the number of drones used. The individual's strengths and weaknesses are assessed, and the best individuals are selected for the next generation. A crossover operation is performed to select parents from the best individuals, and genes are randomly exchanged at crossover points to generate offspring. In this example, individuals are mutated with a probability of μ = 0.1, modifying certain genes in the new individuals to increase population diversity. After 50 iterations, the individuals with the best fitness are selected for path segmentation decoding, ultimately yielding the optimal path plan for each package group that meets the drone payload and flight distance constraints.

[0118] After the lower-level decision module uses the genetic algorithm to output the result, follow the method in step 2.3 to find the state Make an action Rewards , and store it in the memory pool, and extract samples through the priority experience replay mechanism to train the Q network.

[0119] Step 4: Model training, such as Figure 3 The S4 section shows:

[0120] First, initialize the network. In this example, the input state dimension is set to 21 and the output action dimension is set to 16. The target network is initialized to the same parameters as the online network. The Adam optimizer is used with the Step LR scheduler. The memory pool capacity is initialized to N to store the empirical sample data for Q network training. Dueling DQN divides the Q network into a value flow network and an advantage flow network, decomposing the Q value into state value and action advantage. The final action value function is calculated using the following formula:

[0121]

[0122] In the formula Status The value function under In state Take the following The advantage function of To remove the current action The combination of all other action advantages, is the number of actions.

[0123] Packages are assigned sequentially. After obtaining the Q-value of each action through Dueling DQN, an ε-greedy strategy is used to select actions, calculate immediate rewards, and store them in a temporary memory pool. Once all packages have been assigned, a genetic algorithm is used to plan routes by "time-station" grouping. Delayed rewards are calculated and combined with immediate rewards to update the memory pool. A batch of data is sampled from the memory pool by priority, and priorities are updated using the TD error. The network is updated using a weighted mean square error (MSE) loss calculated using importance weights. Parameters are synchronized periodically, and the weights of the Q-network are copied to the target Q-network. Training terminates when the maximum number of iterations is reached, and the model is saved and deployed in the system.

[0124] To demonstrate the superiority of this method over traditional reinforcement learning algorithms for medium- to large-scale problems, experiments were conducted using examples with 100 and 150 parcels, respectively, over a four-hour bus operating period. The delivery destination coordinates for all parcels in this example were randomly generated uniformly within a 15 km x 15 km area. The parcel weights were random integers ranging from 1 to 30 kg, and the delivery time window was a valid time interval within the bus operating period. The parcels were divided into two groups, with a 1:1 ratio: one receiving a 1 yuan delivery reward and the other receiving a 5 yuan delivery reward. The penalty for late delivery was the negative of the delivery reward.

[0125] The present invention selects nine stations on two bus routes as drone stations responsible for drone package delivery in a 15km×15km area. The location coordinates of the nine drone stations are (2000, 2000), (2000, 7000), (2000, 12000), (7000, 2000), (7000, 7000), (7000, 12000), (12000, 2000), (12000, 7000), and (12000, 12000). The drone station with the shortest Euclidean distance to each package is selected as the drop-off point for each package. The maximum weight of packages that can be carried by each bus is 500kg. The flight speed of each drone is 120km / h, the maximum flight distance is 30km, and the maximum payload is 30kg. The drone flight cost per kilometer is 0.08 yuan, and the cost of deploying each drone is 0.05 yuan.

[0126] In order to quantitatively evaluate the performance of different methods on examples of different scales, this paper defines the following indicators:

[0127] Total revenue: from package delivery rewards , Fines for packages not delivered on time , fixed costs of drone deployment and the cost of drone flight distance It consists of four parts, and its formula is:

[0128]

[0129] Delivery rate: The ratio of the number of packages that arrive at the destination on time and correctly within the time window to the total number of packages. The formula is:

[0130]

[0131] The number of packages that arrive at their destination on time and correctly, The total number of packages that need to be delivered.

[0132] Figure 5 and Figure 6 These are the total revenue and delivery rate solved by this method, DQN, and Dueling DQN algorithms when the package demand is 100. Figure 7 and Figure 8The total revenue and delivery rate calculated for this method, the DQN algorithm, and the Dueling DQN algorithm, respectively, when the parcel demand is 150. A comparison of the results obtained by each method shows that this method performs well in terms of total operational revenue and delivery rate, effectively handling medium- to large-scale parcel delivery problems and providing a promising solution for bus-drone combined delivery systems. However, the DQN and Dueling DQN algorithms have limitations when dealing with complex scenarios and cannot meet the requirements for efficient delivery in practical applications.

[0133] This invention integrates idle off-peak bus capacity with drone delivery to create a highly efficient collaborative delivery system combining trunk transport and flexible delivery terminals. A dual-layer decision-making approach, combining a multi-head attention mechanism called Dueling DQN and a genetic algorithm, enables the combined optimization of package-bus matching and drone path planning. This solution not only revitalizes bus resources and improves delivery efficiency, but also reduces traffic congestion and pollution associated with traditional delivery models. It provides a green, sustainable, and innovative solution for smart urban logistics, with practical application value in promoting the coordinated development of urban logistics and public transportation.

Claims

1. A method for urban logistics distribution using public transportation and drones, characterized in that: The specific steps are as follows: Step 1: Construction of urban logistics distribution system and basic data collection. The specific process is as follows: Warehouses will be built at bus terminals on the city's outskirts, integrating warehousing and bus dispatch functions for package collection, distribution, and loading. Bus stops along bus routes located in areas with historically high parcel delivery demand will be upgraded to drone stations. Each station will be equipped with multiple drones and battery swap facilities to support rapid battery replacement upon return. A cargo platform is installed on the top of the bus. When the bus stops at the starting and ending stations, packages are loaded onto the roof before departure. The bus departs according to a fixed departure schedule and travels along a fixed route. When the bus arrives at a regular bus stop, that is, a non-drone stop, it only allows passengers to get on and off. When the bus arrives at a drone stop, not only does it allow passengers to get on and off, but drones also remove packages from the bus that have arrived at their destination for delivery. When the bus stops at the terminal, packages are reloaded onto the bus and the bus continues its journey according to the departure schedule. Obtain information about each bus, drone, and package delivery schedule over the next period. Select a drone station as the destination for each package based on the minimum distance between the package's destination and the drone station. Number each package and select the bus schedule it should take, starting with the smallest number and ending with the largest. Divide the packages into groups based on the drone station they will arrive at and the time they will arrive. Step 2: Model the bus and drone collaborative delivery problem as a Markov decision process. Build an upper-level package-bus schedule matching decision module. Based on the Dueling DQN algorithm with a multi-head attention mechanism, match packages with specific bus schedules. This means selecting the bus that each package should take. The specific process is as follows: Step 2.1: State and action definition The status information is the information of each bus and the information of the package being distributed when distributing the current package. The status information is represented as ,in, Bus feature information, indicating the distribution of packages All The existing load situation of buses on the schedule, Package feature information, which represents the five-dimensional information of the package being distributed, including the number, weight, delivery time window, and rewards and penalties; action Corresponding to the choice of allocating packages to different buses, specifically allocating the current package to One of the bus schedules, Indicates that the package Assigned to bus routes Take a ride; Step 2.2: Multi-head attention mechanism processes state information First, decouple the input state information into dimensional bus feature information and the last 5-dimensional package feature information; two projection layers are used to project the bus and package features into a 128-dimensional embedding space respectively; then the projected bus and package features are concatenated in the sequence dimension to form a sequence of shape [2, batch, 128]; batch is the batch size, that is, the number of experience samples sampled from the memory pool each time; this sequence is input into the multi-head attention layer to capture the association between state information from different subspaces in parallel; After processing by the multi-head attention layer, the output attention result shape is [2, batch, 128]. Then, the attention results of the bus feature information and the package feature information are spliced ​​and converted to the shape of [batch, 256]. The spliced ​​features are input into the subsequent shared layer, and the fusion features are further extracted through the multi-layer fully connected network. Step 2.3: Reward Function Design By building a linkage mechanism between the temporary memory bank and the memory pool, the combination of immediate and delayed rewards is achieved; When the upper layer decides to assign a bus to a package, the immediate reward is calculated on the fly by seeing if the package is assigned to a bus that is overloaded or will not reach the corresponding drone station. , and include the current state ,action , instant rewards and the next state The quadruple Store in temporary memory; When the lower-level decision-making completes the drone path planning through the genetic algorithm, the rewards and penalties for whether each package is delivered on time, the cost of the drone flight distance, and the cost of the drone sorties are added together as the operating income of the group, and the operating income of all groups is calculated. and individual package revenue rewards , divide the two equally and take the weighted sum as the delayed reward ; Extract the corresponding package quad from the temporary memory bank and reward it immediately With delayed rewards The sum of the two is the final reward , and update the quadruple Stored in the memory pool for subsequent model training; reward The calculation formula is as follows: ; ; in, It is the weight parameter used to adjust the individual benefit reward of each package in the delayed reward; is the number of all packages; Step 2.4, Prioritize Experience Replay Introducing a priority experience replay mechanism, which uses a dual mechanism of non-uniform sampling and bias correction to achieve focused learning of high-value experiences; When new experience When stored in the memory pool, its sampling priority is initialized to 1.0 to ensure uniform exploration in the initial stage; as training progresses, the time difference is calculated at each experience replay error: ; in, are the parameters of the current Q network, are the parameters of the target Q network, Indicates the next state possible actions to take; is the current network output, is the target network output; The error plus a small constant of 10 -5 To avoid zero priority, update the priority of the corresponding experience: ; in, To indicate the The probability that an experience is sampled, Indicates the The time difference error of the experience; By exponential transformation Amplify the priority differences to form a non-uniform sampling probability distribution: ; in, To indicate the The probability that an experience is sampled, It is the index of the experience sample, which traverses all the experience samples in the memory pool and is used to calculate the sampling probability When the priority of all experiences is summed, To indicate the Priority of experience of Power, is a hyperparameter used to amplify priority differences; At the same time, in order to offset the distribution bias introduced by priority sampling, the importance sampling weight is introduced : ; in, is the memory pool capacity, is a weight parameter used to correct the deviation caused by non-uniform sampling; Finally, according to the importance sampling weight Calculate the batch average loss by square the time difference error of each sample , and then the gradient is calculated by the back-propagation algorithm, and the optimizer updates the neural network parameters: ; in, is the batch size, that is, the number of experience samples sampled from the experience replay buffer each time; Step 3: Build a lower-level decision-making module to solve the drone delivery scheduling problem of the parcel group based on the genetic algorithm, and implement drone sortie planning, parcel loading combination optimization, and delivery route design for each group. The specific process is as follows: First, the candidate paths are initialized. The target stations are arranged in a random order as the initial solution sequence, which is presented as an integer sequence with a length of X+2, including two 0s and X package destinations. 0 represents the departure and return drone stations, and integers from 1 to X correspond to the delivery package destinations. X is the number of packages that the drone station needs to deliver. The last digit is also 0, indicating that the drone must depart from and return to the drone station. By traversing the intermediate package destination nodes in the complete path represented by the solution sequence, for each node, the corresponding package weight is obtained, a temporary path is constructed that includes the previous path, the node, and the return site, and their distance is calculated. Then, it is checked whether the drone is overloaded after adding this package, and whether the temporary path distance exceeds the maximum flight distance. If any of the conditions are met, the current path plus the return site is stored in the result list as a complete path, and a new path is started; if none of the conditions are met, the node is added to the current path and the load is updated; after the traversal is completed, the return site is added to the last path and stored in the result list, and finally the segmented drone path information is returned; The fitness of each individual is calculated based on the total flight distance after path segmentation and the number of drones used. The individual's quality is evaluated and the better individuals are selected for the next generation. Parents are selected from the better individuals through a crossover operation, and some genes are randomly selected at crossover points to generate offspring. Individuals are mutated with a probability μ, modifying certain genes of the newborn individuals to increase population diversity. After the final iteration, the individual with the best fitness is selected for path segmentation decoding, resulting in a final path planning solution for each package group that meets the drone's payload and flight distance constraints. Step 4: Model training Initialize the Dueling DQN network with a multi-head attention mechanism, set the input state dimension and output action dimension, use the Adam optimizer and Step LR scheduler, and initialize the memory pool capacity. Allocate buses in order of package numbers using the ε-greedy strategy, calculate the immediate reward and store it in a temporary memory bank, and then merge it with the delayed reward and store it in the memory pool. Sample data from the memory pool according to the TD error priority, and update the network by calculating the weighted MSE loss using the importance weights. Regularly synchronize the target network parameters, and save the model after reaching the maximum number of iterations.

2. The urban logistics distribution method using public transportation and drones as claimed in claim 1 is characterized in that: In step 2.2, during the attention calculation process, the sequence is used as 、 and and calculate the attention weights for the interaction between bus and package features; The calculation formula is as follows: ; in, express The dimension size of .

3. The urban logistics distribution method using public transportation and drones as claimed in claim 1, characterized in that: In step 1, the basic data includes: Bus parameters: bus numbers for each bus , maximum load capacity , various bus routes Arrival at each drone station Time ; Drone parameters: Maximum payload , Maximum flight distance , average flight speed ; Package parameters: the number of each package ,weight , destination coordinates , delivery time window , delivery reward Punishment for failure to meet deadlines ; For each package , calculate the Euclidean distance between it and all drone stations, and assign it to the drone station closest to it , forming an ownership relationship; each package must take the bus to the drone station to which it belongs, and then be picked up by the drone for delivery; at the same time Arriving at the same drone station The packages are divided into a package group .