A rider-unmanned vehicle cooperative delivery instant delivery order allocation system

The order allocation system, which combines multi-agent simulation and deep reinforcement learning, solves the problem of untapped heterogeneity between unmanned vehicles and riders, achieves efficient order allocation under dynamic demand, reduces operating costs, and improves customer satisfaction.

CN116415882BActive Publication Date: 2026-04-07TONGJI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-08
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing on-demand delivery order allocation methods fail to effectively utilize the heterogeneity of autonomous vehicles and riders, and struggle to balance optimization effectiveness and computation time under dynamic demands, resulting in insufficient cost and efficiency, especially lacking effective optimization solutions in scenarios involving collaborative delivery of multiple delivery modes.

Method used

An order allocation system based on multi-agent simulation and deep reinforcement learning is adopted. Combining deep reinforcement learning algorithms and maximum utility theory, an order allocation decision model is constructed. The delivery process is simulated through a multi-agent simulation platform to train and iteratively optimize the order allocation strategy, thereby achieving real-time decision support.

Benefits of technology

It improves the accuracy and efficiency of order allocation, reduces operating costs, and enhances customer satisfaction. It is suitable for collaborative delivery scenarios with various delivery modes, especially in collaborative delivery between unmanned vehicles and riders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116415882B_ABST
    Figure CN116415882B_ABST
Patent Text Reader

Abstract

The application provides a rider-unmanned vehicle cooperative distribution instant delivery order distribution system, which comprises an unmanned vehicle and rider cooperative distribution simulation platform based on multiple agents and an order distribution decision system based on deep reinforcement learning and maximum utility theory; the cooperative distribution simulation platform establishes a multiple agent simulation model capable of simulating the conventional process and supply-demand relationship in instant delivery in a region, and the order distribution decision system constructs an order distribution decision model; the order distribution decision model and the multiple agent simulation model interact, train and iterate, and after the training result of the order distribution decision system converges, the real-time instant delivery demand information is input into the cooperative distribution simulation platform, so that real-time decision of order scheduling is realized, and an optimized scheme of distribution order distribution is obtained. The application can help logistics operation personnel to make a relatively optimal order distribution decision under dynamic continuous demand, realize the transformation of the intelligent city distribution system and the cost reduction and benefit increase of logistics enterprises.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of logistics order allocation optimization technology, and relates to an order allocation optimization system (computer intelligent calculation and application) that considers the heterogeneity of riders and unmanned vehicles and the future revenue of order delivery in the context of instant delivery. Background Technology

[0002] In recent years, with the rapid development of online shopping, on-demand delivery services have gradually become an important part of urban logistics. On-demand delivery services deliver goods requested by consumers online within two hours via delivery personnel; delivery time is a crucial factor determining service quality. Typical on-demand delivery items include fresh produce, takeout, medicine, and urgent documents. The on-demand delivery industry has maintained a high growth rate in recent years, and to ensure good timeliness even with surging demand, a large amount of manpower has been deployed to delivery services. Therefore, labor costs account for more than half of the total operating costs for on-demand delivery. Driverless vehicles, due to their easier control, less weather-dependent delivery capabilities, and lack of need for rest or shifts, are widely regarded as a future alternative to manual delivery and have begun pilot operations in many countries, including China, the United States, the United Kingdom, Switzerland, Japan, and Estonia. However, since driverless vehicles cannot effectively complete tasks such as order delivery in elevator-less buildings and communication and negotiation with customers in return and exchange scenarios, and the government has not yet introduced mature laws and regulations on the supervision of driverless vehicles and accident liability, in the foreseeable future, such as 2023-2028, driverless vehicles and riders will serve as a transitional state to jointly serve the order delivery of instant delivery scenarios.

[0003] Currently, there are no readily available methods or optimized routes for order allocation between autonomous vehicles and delivery riders. Using methods from traditional manual delivery ignores the heterogeneity between autonomous vehicles and riders, failing to fully utilize the potential of autonomous vehicles. Furthermore, existing heuristic-based order allocation methods rarely balance optimization effectiveness and computation time under large-scale dynamic demands, leading to cost and efficiency deficiencies. For example, patent CN109598366A proposes a scheduling optimization method for food delivery, based on an improved ant colony algorithm and a variable neighborhood local search strategy, combined with an adaptive parameter correction mechanism, to guide the algorithm to perform effective global search and improve local optimization capabilities, but it does not consider the dynamic nature of orders and the heterogeneity of delivery personnel. Patent CN114970103A considers the randomness of travel time caused by delivery personnel experience and actual delivery conditions, combining simulation methods and heuristic algorithms to solve the order allocation and route planning for delivery, but it does not consider dynamically added orders in real time.

[0004] Most existing instant delivery order allocation methods are based on problem scenarios with only one delivery mode. They use heuristic or precise algorithms to optimize delivery resource scheduling when only considering the current order. They cannot be directly applied to scenarios with multiple delivery modes coordinating delivery. Furthermore, the optimization results may limit the optimization potential for future orders. At the same time, heuristic algorithms often require a trade-off between computation time and optimization effect. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of prior art and disclose an order allocation system suitable for collaborative delivery by unmanned vehicles and riders. This system considers both the future benefits of immediate delivery order allocation and the heterogeneity of key parameters of different delivery methods. It can obtain order allocation results in real time based on current information, thereby helping logistics operators make better order allocation decisions under dynamic and continuous demand, realizing the transformation of smart city delivery systems and cost reduction and efficiency improvement for logistics companies.

[0006] Technical solution:

[0007] An instant delivery order allocation system for rider-unmanned vehicle collaborative delivery includes a multi-agent unmanned vehicle and rider collaborative delivery simulation platform and an order allocation decision system based on deep reinforcement learning and maximum utility theory. The logical relationship is as follows: the deep reinforcement learning algorithm used in the order allocation decision system requires a policy training and evaluation environment, and the feature parameters for utility evaluation based on maximum utility theory are both derived from the multi-agent unmanned vehicle and rider collaborative delivery simulation platform.

[0008] Furthermore, the collaborative delivery simulation platform establishes a multi-agent simulation model capable of simulating the routine processes and supply-demand relationships in instant delivery within a region; the order allocation decision system constructs an order allocation decision model, comprising two parts: delivery mode selection based on deep reinforcement learning and specific delivery vehicle allocation based on maximum utility theory. The order allocation decision model interacts, trains, and iterates with the multi-agent simulation model. After the training results of the order allocation decision system converge, inputting real-time instant delivery demand information into the collaborative delivery simulation platform enables real-time decision-making on order scheduling, resulting in an optimized solution for delivery order allocation.

[0009] The order allocation decision-making method proposed in this invention, after being constructed and trained in a simulation model, can continue to train and iterate in actual use as new order information and decision feedback records are added. This allows for the establishment of a neural network that is more targeted to the application scenario and more accurate in predicting decision-making effects, thus obtaining real-time order matching decisions. Therefore, the method proposed in this invention can theoretically be applied to order allocation for collaborative delivery of unmanned vehicles and riders in all regions of on-demand delivery.

[0010] Compared with the prior art, the present invention has the following beneficial effects:

[0011] 1. This invention proposes an order allocation system suitable for collaborative delivery between unmanned vehicles and riders. It can scientifically reflect the operation process of instant delivery and is applicable to the actual needs of operational decision-making. It can help logistics suppliers evaluate and improve the operation and management strategies of unmanned vehicle delivery. At the same time, the model also has the potential to be applied to research in other fields, especially order allocation decisions with two or more service modes, such as order allocation decisions for human-driven taxis and unmanned taxis under the same operating platform.

[0012] 2. This invention proposes a simulation-based optimization decision-making scheme that combines multi-agent simulation and deep reinforcement learning. It creates an artificial dynamic environment for decision-makers, sets different interaction and reward rules, observes the related behavioral outcomes, and ultimately provides decision support for real-time order allocation. This invention improves the theoretical system of order allocation in dynamic environments and provides a solution to the problem of order allocation in collaborative delivery between riders and unmanned vehicles, which can facilitate the large-scale application of unmanned vehicle delivery.

[0013] This invention provides a technical solution for order allocation in collaborative delivery between riders and unmanned vehicles. It can also be applied to order allocation decisions for different operating modes under the same operating system, such as order allocation decisions for human-driven taxis and unmanned taxis under the same operating platform. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the system relationships of the present invention.

[0015] Figure 2 It is the overall process of instant delivery of orders through collaboration between unmanned vehicles and riders. Detailed Implementation

[0016] The technical solutions provided in this application will be further described below with reference to specific embodiments and accompanying drawings. The advantages and features of this application will become clearer from the following description.

[0017] It should be noted that the embodiments of this application are preferred for implementation and are not intended to limit the application in any way. The technical features or combinations of technical features described in the embodiments of this application should not be considered isolated; they can be combined with each other to achieve better technical effects. The scope of the preferred embodiments of this application may also include other implementations, and this should be understood by those skilled in the art to which the embodiments of this application pertain.

[0018] Techniques, methods, and apparatus known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and apparatus should be considered part of the specification. In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limiting. Therefore, other examples of exemplary embodiments may have different values.

[0019] The accompanying drawings in this application are all in a very simplified form and use non-precise proportions, intended only to facilitate and clarify the illustration of the embodiments of this application, and are not intended to limit the implementation of this application. Any modifications to the structure, changes in the proportional relationships, or adjustments to the size, without affecting the effects and purposes achieved by this application, should fall within the scope of the technical content disclosed in this application. Furthermore, the same reference numerals appearing in the various drawings of this application represent the same features or components, and can be applied to different embodiments.

[0020] The on-demand delivery described in this application primarily targets goods such as fresh produce, pharmaceuticals, and supermarkets, which originate from centralized distribution stations. This application example focuses on an 8km × 5km fresh produce delivery service area, with an average of 1500 orders requiring delivery daily. Customer order times are shown in Table 1. Late-night orders (22:00-6:00) account for 8% of total orders, with peak order periods concentrated between 10:30-13:00 and 20:00-21:00, accounting for 24% and 15% of total orders respectively. The maximum acceptable delivery time for customers is shown in Table 2. Orders within 30 minutes account for 6%, within 31-60 minutes for 44%, within 61-90 minutes for 27%, within 91-210 minutes for 17%, and over 211 minutes for 6%.

[0021] Table 1. Customer Order Placement Time Distribution

[0022]

[0023]

[0024] Table 2 Maximum acceptable delivery time for customers

[0025]

[0026] like Figure 1As shown, this example demonstrates an order allocation system for collaborative delivery between unmanned vehicles and riders in instant delivery, including a simulation platform for collaborative delivery between unmanned vehicles and riders based on multi-agent intelligence and an order allocation decision system based on deep reinforcement learning and maximum utility theory. The logical relationship is as follows: the deep reinforcement learning algorithm used in the order allocation decision system requires a policy training and evaluation environment, and the feature parameters for utility evaluation based on maximum utility theory are both derived from the simulation platform for collaborative delivery between unmanned vehicles and riders based on multi-agent intelligence.

[0027] Furthermore, the logical relationship between the two is as follows: the collaborative delivery simulation platform is a multi-agent simulation model that can simulate the routine process and supply and demand relationship in instant delivery within a region; the order allocation decision system is an order allocation decision model that includes two parts: delivery mode selection based on deep reinforcement learning and specific delivery vehicle allocation based on maximum utility theory. The order allocation decision model interacts, trains, and iterates with the multi-agent simulation model. After the training results of the order allocation decision system converge, inputting real-time instant delivery demand information into the collaborative delivery simulation platform can realize the real-time decision-making of order scheduling by the system of the present invention and obtain an optimized solution for delivery order allocation.

[0028] Furthermore, the collaborative delivery simulation platform includes a multi-agent simulation model, specifically customer agent modeling, rider / unmanned vehicle agent modeling, and delivery station agent modeling, as well as a simulation environment building module. The simulation environment assumes the following conditions: (1) all unmanned vehicles and riders depart from the delivery station and return to the delivery station; (2) all customers can only be served by unmanned vehicles or riders; (3) all riders or unmanned vehicles have the same onboard capacity and battery capacity; (4) all riders or unmanned vehicles deliver at least one order per trip; (5) riders work from 6:00 to 22:00, and unmanned vehicles work 24 hours a day; (6) the delivery station has all the goods needed by the customer, does not need to transfer goods from other delivery stations, and customers will not be rejected due to insufficient inventory.

[0029] The agent-based modeling in the collaborative delivery simulation platform is used to simulate the routine processes in instant delivery and the behaviors and interactions of e-commerce platforms, riders or autonomous vehicles, and customers, thereby creating a training and evaluation environment for the order allocation algorithm. The agent-based model constructed in this invention can initialize the distribution of delivery demand and supply based on order data obtained from actual operations, delivery capacity parameters of delivery personnel and autonomous vehicles, geographical information of the delivery area, and road network, simulating the real delivery process to allocate orders and schedule riders and autonomous vehicles. Three types of agents are constructed in the agent-based model to simulate the relevant stakeholders in instant delivery: customers, riders or autonomous vehicles, and delivery stations. These three types of agents are defined below.

[0030] The customer agent modeling includes customer attributes such as location, order placement time, maximum acceptable delivery time, and satisfaction level. The main customer behaviors include submitting the order at the time of placement and calculating their satisfaction upon delivery completion. Satisfaction is a non-negative value (not greater than 1) reflecting the customer's assessment of delivery completion based on the delivery time. The specific calculation method is as follows:

[0031]

[0032] Where D i and T i These represent the maximum acceptable delivery time and order placement time for customer i, respectively. ji The time it takes for a rider or autonomous vehicle j to complete the delivery of customer i's order. If the order cannot be successfully delivered, the satisfaction level is set to 0.

[0033] The rider / autonomous vehicle intelligent agent modeling defines riders and autonomous vehicles as the same type of intelligent agent with identical behavioral logic but different attribute values. These attributes include operating cost per kilometer, current location, speed, maximum capacity, delivery task list, remaining range, and working time period. Key behaviors include pickup, route planning, battery replacement, movement, and recording movement trajectory and distance. Each time the rider returns to the delivery station to pick up goods and after completing an order, both the rider and autonomous vehicle check the remaining range. If the remaining range is less than 10 kilometers, they return to the delivery station to replace the battery and reset the range. Specific operational parameters for the autonomous vehicle and rider are shown in Table 3. The rider's scheduling plan is shown in Table 4.

[0034] Table 3 Specific parameters for unmanned vehicle and rider operation

[0035] project driverless car rider Maximum driving speed (km / h) 40 40 Maximum vehicle capacity (orders) 15 10 Battery replacement time (s) 60 60 Battery range (km) 100 60 Working hours 24 / 7 6:00-22:00 Operating cost per kilometer (RMB) 1.0 2.0

[0036] Table 4 Rider Scheduling Plan

[0037] Working hours Number of working riders 6:00-7:00 4 7:00-20:00 25 20:00-22:00 5

[0038] The delivery station intelligent agent model defines a delivery station as an intelligent agent that decides how to deliver customer orders (by rider or autonomous vehicle) and which specific rider or autonomous vehicle will be responsible for delivery. It also serves as the location for riders / autonomous vehicles to pick up orders and replace batteries, and as the origin and destination of deliveries. Its attributes include location and a list of orders to be assigned. Its main behaviors include deciding on the delivery method for orders, selecting a specific delivery rider or autonomous vehicle, and updating the list of orders to be assigned.

[0039] The simulation environment building module builds a simulation environment that includes geographical information such as roads and residences, as well as time information. The geographical information is imported from a shapefile file, and the time information is updated every 1 minute.

[0040] In the simulation environment, the interactive behaviors of various intelligent agents include the following defined events: customer order placement, order allocation, rider or autonomous vehicle pickup, rider or autonomous vehicle movement, rider or autonomous vehicle battery replacement, customer receipt of goods, etc. The specific content of each event and its order of occurrence on the timeline are as follows (e.g., Figure 2 As shown):

[0041] When the S1 environment time reaches the customer's order time, the customer sends an order signal to the delivery station, and the delivery station puts the corresponding order into the list of orders to be assigned.

[0042] S2 extracts the order features, spatiotemporal distribution features of riders and unmanned vehicles, remaining capacity of riders and unmanned vehicles, and estimated return time of riders and unmanned vehicles to the delivery station for each order in the order list to be assigned. Based on this, it first makes a decision on the delivery mode (rider or unmanned vehicle) of the order based on deep reinforcement learning, and then makes a decision on the specific rider or unmanned vehicle for delivery based on the utility maximization theory.

[0043] S3 riders and autonomous vehicles pick up goods at the delivery station and plan delivery routes based on accepted orders. They then depart from the delivery station and proceed along the road network to deliver the goods to the customers one by one.

[0044] During operation, the S4 rider and autonomous vehicle continuously monitor their remaining range. If the remaining range is insufficient to complete the current order, proceed to the next customer's location, and then return to the delivery station, the delivery route is modified to allow the rider or autonomous vehicle to return to the delivery station to replace the battery after completing the current order, and to pick up any orders already assigned to that rider or autonomous vehicle. Similarly, if the rider or autonomous vehicle returns to the delivery station empty after completing all orders, and the remaining range is less than 10 kilometers, a battery replacement request is also made, and any orders already assigned to that rider or autonomous vehicle are picked up.

[0045] After the S5 rider or autonomous vehicle arrives at the customer's location, the customer picks up the goods, and the delivery time is calculated and satisfaction is assessed.

[0046] The constructed and trained order allocation decision system includes a feature extraction module, a delivery mode selection module based on deep reinforcement learning, and a specific delivery vehicle allocation module based on maximum utility theory.

[0047] The delivery mode selection module based on deep reinforcement learning is implemented based on the deep Q-network algorithm. It extracts order features, rider and autonomous vehicle spatiotemporal distribution features, rider and autonomous vehicle remaining capacity, and rider and autonomous vehicle expected return time to the delivery station from each order in the order list to be assigned in the simulation environment through the feature extraction layer (i.e. Q-network input layer). These features are then input into the delivery mode selection module based on deep reinforcement learning.

[0048] The delivery mode selection module based on deep reinforcement learning fully considers the continuity and dynamism of orders in instant delivery, runs the order allocation decision in the scenario of autonomous vehicle and rider collaborative delivery, and determines which delivery mode to use for each order, thereby maximizing future revenue;

[0049] The specific delivery vehicle allocation module based on the maximum utility theory determines which specific unmanned vehicle or rider will be used for delivery of each order.

[0050] Furthermore, the delivery mode selection module based on deep reinforcement learning specifically includes a Markov decision model and a deep reinforcement learning training algorithm, wherein:

[0051] In the modeling of the Markov decision-making model, this invention models the problem as a Markov decision process when determining the delivery mode. This process enables the agent to maximize long-term cumulative benefits through reasonable decisions, thereby achieving optimization of long-term returns. The various parts of this process in this invention are defined as follows:

[0052] (1) State: State s t This section describes a problem scenario at a specific decision-making moment and consists of three parts: the decision-making time, demand information, and supply information. Demand information refers to the information about the orders to be assigned, including their location, maximum acceptable delivery time, and order placement time. Supply information is the information about the riders or autonomous vehicles needed for the decision, including their location, remaining capacity, and estimated return time to the delivery station.

[0053] (2) Action: Action a t The final decision is made by the delivery station's intelligent agent. This invention includes two actions: assigning an order to rider delivery and assigning an order to unmanned vehicle delivery. Specifically, this involves adding an order to either the rider delivery order list or the unmanned vehicle delivery order list.

[0054] (3) Reward: Reward R t Is the agent in state s t The following action a was taken t The rewards earned in the present will be distributed in the future. After the allocation method is determined and the delivery riders or autonomous vehicles are selected, the rewards calculated based on the estimated delivery time and operating costs will be fed back to the delivery station's intelligent agent.

[0055] (4) State Transition: Once an order is added to the order list for a specific delivery method, it is removed from the delivery station's list of orders awaiting delivery. The rider's and the autonomous vehicle's location, remaining capacity, and estimated return time are also updated accordingly, and the state s... t Transition to the next state s t ′.

[0056] (5) Discount factor: The discount factor γ is used to calculate the present value of future rewards in order to balance the emphasis on future gains and present gains.

[0057] The deep reinforcement learning training algorithm specifically uses the deep Q-network algorithm for solving:

[0058] The deep Q-network algorithm, based on the above definition in the Markov decision model, can be derived using formula (2) in state s. t Take action a t The long-term cumulative return Q(s) t a t ), and thus, given Q(s) t a t In the case of Q(s), the action that maximizes long-term cumulative return is selected to achieve decision-making that considers long-term benefits. t a t The prediction of the Q network is achieved by using a backpropagation deep neural network to extract state and action features. The training of the model, i.e. the update of the Q network, is mainly achieved by minimizing the mean square error between the Q estimate and the Q target value generated by the Q network and the target network with the same initial parameters, respectively. The objective function of this minimization is shown in formula (3).

[0059]

[0060]

[0061] Where r represents the immediate reward for taking action a in state s. and θ i These are the parameters of the target Q-network and the Q-network, respectively. They are identical during initialization, and θ is adjusted periodically during subsequent training. i Copy the value to middle.

[0062] The detailed process of the deep Q-network algorithm is shown in the following pseudocode:

[0063]

[0064]

[0065] The specific delivery vehicle allocation module based on the maximum utility theory: After determining the delivery mode through deep reinforcement learning training algorithm, the present invention will decide on the specific vehicle (a rider or unmanned vehicle) to deliver the order based on the maximum utility theory. That is, for each order, the utility value of several riders or several unmanned vehicles to deliver the order is calculated one by one, and the rider or unmanned vehicle with the highest utility value is selected to perform the delivery task.

[0066] Based on the current characteristic parameters of the rider or unmanned vehicle and the customer, the calculation method of the order allocation utility value is shown in formula (4).

[0067]

[0068] in d represents the time that customer i has waited until the decision-making time. ij The increased driving distance for riders or autonomous vehicles due to delivering customer orders, v j D represents the average speed of the rider or autonomous vehicle j. rj For the remaining range of the rider or autonomous vehicle j, d o The distance still needed for riders or autonomous vehicles to deliver orders that have not yet been completed. w α e α r These are three weighting coefficients ranging from 0 to 1. The calculation of this utility value prioritizes orders that have been waiting for a longer period for delivery by riders or autonomous vehicles with longer remaining range and shorter additional delivery distances after accepting the order. This specific vehicle allocation based on maximum utility theory allows order allocation decisions to better consider the differences between supply and demand.

[0069] The simulation is run based on order data from the region over a period of time, allowing the Q-network in the allocation decision model to learn and train in the simulation environment until the sum of customer satisfaction for all orders in the simulation converges. After that, the allocation decision model is connected to the actual order allocation system to make real-time decisions on the allocation of actual orders.

[0070] Because autonomous vehicles and riders collaborate in delivery, the depth of autonomous vehicle involvement is closely related to the development of autonomous driving technology and the laws and regulations governing autonomous vehicle delivery operations. Therefore, by changing the ratio of autonomous vehicles to riders, eight scenarios were constructed, including pure rider delivery and autonomous vehicle and rider collaborative delivery, as shown in Table 5.

[0071] Table 5 Number of driverless vehicle and rider fleets in different scenarios

[0072] Scene Number Number of driverless cars Number of riders 1-23 1 23 2-20 2 20 3-18 3 18 4-16 4 16 5-14 5 14 8-8 8 8 12-2 12 2 BAU 0 25

[0073] For the delivery of 1500 orders per day in this region in this example, the allocation strategy based on deep reinforcement learning and maximum utility theory constructed in this invention is compared with the allocation strategy based on greedy algorithm and the allocation strategy based on KM algorithm in terms of total customer satisfaction and total operating cost. The results are shown in Tables 6 and 7:

[0074] Table 6 Overall Customer Satisfaction

[0075]

[0076]

[0077] Table 7 Total Operating Costs

[0078]

[0079] Compared to traditional allocation strategies based on greedy algorithms and KM algorithms, the method of this invention can improve total customer satisfaction by up to 7.62% and reduce total operating costs by 60.33% in this instance, with the same combination of riders and unmanned vehicles. This result shows that the order allocation decision method based on deep reinforcement learning and maximum utility theory described in this invention can improve the existing level of on-demand delivery services in terms of both service level and operating costs.

[0080] The above description is merely a description of preferred embodiments of this application and is not intended to limit the scope of this application in any way. Any changes or modifications made by those skilled in the art based on the above-disclosed technical content should be considered as equivalent and valid embodiments and fall within the scope of protection of the technical solution of this application.

Claims

1. A rider-unmanned vehicle collaborative delivery instant delivery order allocation system, characterized in that, The system includes a multi-agent-based unmanned vehicle and rider collaborative delivery simulation platform and an order allocation decision system based on deep reinforcement learning and maximum utility theory. The logical relationship is as follows: the deep reinforcement learning algorithm used in the order allocation decision system requires a policy training and evaluation environment, and the feature parameters for utility evaluation based on maximum utility theory are both derived from the multi-agent-based unmanned vehicle and rider collaborative delivery simulation platform. The collaborative delivery simulation platform is a multi-agent simulation model that can simulate the routine processes and supply and demand relationships in instant delivery within a region. The order allocation decision system, namely the construction of an order allocation decision model, includes two parts: delivery mode selection based on deep reinforcement learning and specific delivery vehicle allocation based on maximum utility theory. The order allocation decision model interacts, trains, and iterates with the multi-agent simulation model. After the training results of the order allocation decision system converge, real-time delivery demand information is input into the collaborative delivery simulation platform to realize real-time decision-making on order scheduling and obtain an optimized solution for delivery order allocation. The collaborative delivery simulation platform includes a multi-agent simulation model, specifically customer agent modeling, rider / unmanned vehicle agent modeling, delivery station agent modeling, and also includes a simulation environment building module; the simulation environment assumes the following conditions: (1) all unmanned vehicles and riders depart from the delivery station and return to the delivery station; (2) all customers can only be served by unmanned vehicles or riders; (3) all riders or unmanned vehicles have the same vehicle capacity and battery capacity; (4) all riders or unmanned vehicles deliver at least one order per trip; (5) riders work from 6:00 to 22:00, and unmanned vehicles work 24 hours a day; (6) the delivery station has all the goods needed by the customer, does not need to transfer goods from other delivery stations, and the customer will not be rejected due to insufficient inventory. The customer agent modeling includes customer attributes such as location, order placement time, maximum acceptable delivery time, and satisfaction level. The main customer behaviors include submitting the order at the time of placement and calculating their satisfaction upon delivery completion. Satisfaction is the customer's assessment of delivery completion based on the delivery time, and is a non-negative value not greater than 1. The specific calculation method is as follows: in and For customers The maximum acceptable delivery time and order time. For riders or autonomous vehicles Complete customer The order delivery time; if the order cannot be successfully delivered, the satisfaction level is 0. The rider / autonomous vehicle intelligent agent modeling describes riders and autonomous vehicles as the same type of intelligent agents with the same behavioral logic but different attribute values. Their attributes include operating cost per kilometer, current location, speed, maximum capacity, delivery task list, remaining range, and working time period. Their main behaviors include picking up goods, planning delivery routes, changing batteries, moving, and recording movement trajectory and distance. Each time the rider returns to the delivery station to pick up goods and after completing an order, the rider and autonomous vehicle will check the remaining range. If it is less than 10 kilometers, they will return to the delivery station to change the battery and reset the range. The delivery station intelligent agent model is a smart agent that decides how to deliver customer orders and which rider or autonomous vehicle will deliver them. It is also the place where riders / autonomous vehicles pick up goods and replace batteries, as well as the origin and destination of delivery. Its attributes include location and a list of orders to be assigned. Its main behaviors include deciding on the delivery method of the order, selecting a specific delivery rider or autonomous vehicle, and updating the list of orders to be assigned. The simulation environment building module builds a simulation environment that includes geographical information such as roads and residences, as well as time information. The geographical information is imported from a shapefile file, and the time information is updated every 1 minute. In the simulation environment, the interactive behaviors of various intelligent agents include the following defined events: customer order placement, order allocation, rider or autonomous vehicle pickup, rider or autonomous vehicle movement, rider or autonomous vehicle battery replacement, customer receipt of goods, etc.; the specific content of each event and its order of occurrence on the timeline are as follows: When the S1 environmental time reaches the customer's order time, the customer sends an order signal to the delivery station, and the delivery station puts the corresponding order into the list of orders to be assigned. S2 extracts the order features, spatiotemporal distribution features of riders and unmanned vehicles, remaining capacity of riders and unmanned vehicles, and estimated return time of riders and unmanned vehicles to the delivery station for each order in the order list to be assigned. Based on this, it first makes a decision on the order delivery mode based on deep reinforcement learning, and then makes a decision on the specific rider or unmanned vehicle for delivery based on the utility maximization theory. S3 riders and autonomous vehicles pick up goods at the delivery station and plan delivery routes based on accepted orders. They then depart from the delivery station and travel along the road network to deliver the goods to the customers one by one. The S4 continuously monitors the remaining range of riders and autonomous vehicles during their movement. If the remaining range is insufficient to deliver the current order, proceed to the next customer's location, and then return to the delivery station, the delivery route is modified to allow the rider or autonomous vehicle to return to the delivery station to replace the battery after delivering the current order, and to pick up the orders already assigned to that rider or autonomous vehicle. After delivering all orders and returning to the delivery station empty, if the remaining range is less than 10 kilometers, the rider or autonomous vehicle also requests a battery replacement and picks up the orders already assigned to that rider or autonomous vehicle. After the S5 rider or autonomous vehicle arrives at the customer's location, the customer picks up the goods, and the delivery time is calculated and satisfaction is assessed.

2. The rider-unmanned vehicle collaborative delivery instant delivery order allocation system as described in claim 1, characterized in that, A decision-making system for order allocation is constructed and trained, including a feature extraction module, a delivery mode selection module based on deep reinforcement learning, and a specific delivery vehicle allocation module based on maximum utility theory. The delivery mode selection module based on deep reinforcement learning is implemented using a deep Q-network algorithm. The feature extraction layer, i.e. the Q-network input layer, extracts the order features, rider and autonomous vehicle spatiotemporal distribution features, rider and autonomous vehicle remaining capacity, and rider and autonomous vehicle expected return time to the delivery station from each order in the list of orders to be assigned in the simulation environment. These features are then input into the delivery mode selection module based on deep reinforcement learning. The delivery mode selection module based on deep reinforcement learning fully considers the continuity and dynamism of orders in instant delivery, runs the order allocation decision in the scenario of autonomous vehicle and rider collaborative delivery, and determines which delivery mode to use for each order, thereby maximizing future revenue; The specific delivery vehicle allocation module based on the maximum utility theory determines which specific unmanned vehicle or rider will be used for delivery of each order.

3. The rider-unmanned vehicle collaborative delivery instant delivery order allocation system as described in claim 2, characterized in that, The delivery mode selection module based on deep reinforcement learning specifically includes a Markov decision model and a deep reinforcement learning training algorithm, wherein: The Markov decision model is used to model the problem of delivery mode as a Markov decision process. This process enables the agent to maximize long-term cumulative benefits through reasonable decisions, thereby achieving optimization of long-term benefits. The components of this process are defined as follows: (1) State: State It is used to describe a problem scenario at a certain decision point and consists of three parts: decision time point, demand information, and supply information. Demand information is the information of the orders to be allocated, including location, maximum acceptable delivery time and order time. Supply information is the information of the riders or unmanned vehicles required for the decision, including location, remaining capacity and estimated time to return to the delivery station for pickup. (2) Action: Action The final decision is made by the delivery station's intelligent agent; two actions are set up, namely setting the order as rider delivery and setting the order as unmanned vehicle delivery. In specific operation, this is reflected in adding a certain order to the order list assigned to rider delivery or the order list assigned to unmanned vehicle delivery. (3) Rewards: Rewards Is the agent in a state? The following actions were taken. The revenue gained in the present and future; after the allocation method is determined and the delivery riders or unmanned vehicles are selected, the rewards calculated based on the estimated delivery time and operating costs will be fed back to the delivery station's intelligent agent; (4) State Transition: Once an order is added to the order list for a certain delivery method, it is removed from the delivery station's list of orders to be delivered. The rider's and the autonomous vehicle's location, remaining capacity, and estimated return time are also updated accordingly. Transition to the next state ; (5) Discount factor: Discount factor Used to calculate the present value of future rewards in order to balance the emphasis on future gains and present gains; The deep reinforcement learning training algorithm specifically uses the deep Q-network algorithm for solving: The deep Q-network algorithm, based on the above definition in the Markov decision model, can be derived in state (2) according to formula (2). Take action below Long-term cumulative returns Therefore, it is possible to know In situations where the action that maximizes long-term cumulative returns is chosen, decision-making should consider long-term benefits; among which... The prediction is achieved by extracting state and action features through backpropagation deep neural network. The training of the model, i.e. the update of the Q network, is achieved by minimizing the mean square error between the Q estimate and the Q target value generated by the Q network and the target network with the same initial parameters, respectively. The objective function of this minimization is shown in Equation (3). (2) (3) in Represents the state Take action below Instant rewards and These are the parameters of the target Q-network and the Q-network, respectively. They are identical during initialization and are adjusted periodically during subsequent training. Copy the value to middle.

4. The rider-unmanned vehicle collaborative delivery instant delivery order allocation system as described in claim 2, characterized in that, The specific delivery vehicle allocation module based on the maximum utility theory: After determining the delivery mode through deep reinforcement learning training algorithm, it will decide on the specific vehicle to deliver the order based on the maximum utility theory. That is, for each order, it will calculate the utility value of several riders or several unmanned vehicles to deliver the order and select the rider or unmanned vehicle with the highest utility value to execute the delivery task. Based on the current characteristic parameters of the rider or unmanned vehicle and the customer, the calculation method of the order allocation utility value is shown in formula (4); (4) in For customers The time already spent waiting up to the decision-making moment For riders or autonomous vehicles Due to delivery customers The increased driving distance due to orders, For riders or autonomous vehicles average driving speed For riders or autonomous vehicles The remaining driving range, For riders or autonomous vehicles The distance still needs to be traveled for orders that have not yet been delivered; , , These are three weighting coefficients with values ​​between 0 and 1.

Citation Information

Patent Citations

  • An optimal scheduling method for a take-out delivery process

    CN109598366A

  • Model training method and device and unmanned equipment scheduling method and device

    CN113298445A

  • Article distribution scheduling method and device, computer equipment and storage medium

    CN113570312A