A dynamic dispatch optimization method for delivery vehicles based on DDQN algorithm

By applying a dynamic scheduling optimization method based on the DDQN algorithm in the scheduling of fresh food delivery vehicles, the problems of low scheduling efficiency and reduced timeliness in the prior art are solved, and more efficient resource utilization and order satisfaction are achieved.

CN117726040BActive Publication Date: 2025-05-06ANHUI AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311830634.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-05-06
Estimated Expiration
2043-12-27

AI Technical Summary

Technical Problem

In the scheduling of fresh food distribution vehicle, the combination complexity of the allocation space is too high, resulting in low system resource utilization and scheduling efficiency, which in turn leads to the problem of decreasing the timeliness of fresh food products.

Method used

The dynamic scheduling optimization method based on the DDQN algorithm is adopted to treat the scheduling problem of fresh food delivery vehicles as a continuous time process based on the SMDP framework, simulated through a discrete event simulator, and trained dual agents to optimize the scheduling decisions under "new order events" and "vehicle events".

Benefits of technology

It significantly reduces the combination complexity of the allocation space, improves the system resource utilization and scheduling efficiency, ensures the timeliness of fresh products and the timely delivery of orders, and meets the diversified needs of orders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117726040B_ABST
    Figure CN117726040B_ABST
Patent Text Reader

Abstract

The present invention discloses a dynamic dispatch optimization method for delivery vehicles based on the DDQN algorithm, and belongs to the technical field of fresh food delivery vehicle dispatch based on deep reinforcement learning; the present invention regards the dynamic vehicle dispatch problem of fresh food delivery as a continuous time process, models it based on the SMDP (Semi-Markov Decision Process) framework, and uses the DDQN (Double Deep Q-Learning) algorithm to train dual agents to make dispatch allocations when processing "new order events" and "vehicle events". This method significantly reduces the combinatorial complexity of the allocation space, and shows better average allocation time while considering multiple allocation constraints. By improving system resource utilization and dispatch efficiency, the problem of decreased timeliness of fresh products due to delays in fresh food delivery is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fresh food delivery vehicle scheduling, and in particular to a delivery vehicle dynamic scheduling optimization method based on a DDQN algorithm. Background Art

[0002] In recent years, the market demand for fresh products such as meat, fruits, and vegetables in my country has grown rapidly. Due to the perishability and shelf life of fresh products, timeliness requirements are very high. Ensuring timely delivery is crucial to maintaining the freshness and quality of products. Delivery vehicles must deliver products from transit stations or warehouses in the supply chain to their destinations in the shortest possible time to ensure the freshness and quality of products. Orders for fresh products are usually diverse, and various products have different characteristics and processing requirements. For example, they need to be refrigerated, frozen, or maintain a specific humidity. The special needs of different types of products need to be considered when scheduling vehicles. Deep reinforcement learning technology needs to be used to improve the efficiency of dynamic vehicle scheduling for fresh food delivery.

[0003] In the related art, when studying the dispatch of fresh food delivery vehicles, this type of problem is often regarded as an MDP (Markov Decision Process), assuming that the allocation occurs within a fixed time interval; since all possibilities must be considered during the allocation, the many-to-many allocation scheduling problem will lead to a high combinatorial complexity of the allocation space, which will not only reduce the system resource utilization and scheduling efficiency, but also simplify the allocation restriction factors, which will easily cause delays in fresh food delivery and reduce the timeliness of fresh food products. In order to solve the above problems, the present invention proposes a dynamic dispatch optimization method for delivery vehicles based on the DDQN algorithm. Summary of the invention

[0004] 1. Technical issues to be solved

[0005] The purpose of the present invention is to propose a dynamic vehicle scheduling optimization method for fresh food delivery based on the DDQN algorithm to solve the problems raised in the background technology. The present invention significantly reduces the combinatorial complexity of the allocation space, and shows a better average allocation time while considering multiple allocation constraints. By improving the utilization rate of system resources and scheduling efficiency, the problem of reduced timeliness of fresh products caused by delays in fresh food delivery is solved.

[0006] (II) Technical solution

[0007] In order to achieve the above object, the technical solution adopted by the present invention is:

[0008] A distribution vehicle dynamic scheduling optimization method based on DDQN algorithm includes the following steps:

[0009] S1. The dynamic vehicle scheduling problem in fresh food delivery is regarded as a continuous time process based on the SMDP (Semi-Markov Decision Process) framework: This method proposes an event-based SMDP formula based on the characteristics that fresh food delivery orders appear randomly over time and the time intervals between continuous allocations are random, and defines the basic components of SMDP: environment, state, action space, reward function and environmental dynamics. In the system, two important events that trigger allocation are clearly defined: "new order event" and "vehicle event", which simplifies the original many-to-many allocation scheduling problem into a one-to-many allocation scheduling problem.

[0010] S2. Simulation using Discrete Event Simulation (DES): This method uses Python to configure the simulator. The simulator consists of three types of objects: basic entities, agents, and environments. The simulator works by maintaining a list of orders in chronological order and using specific processing routines to handle these events.

[0011] During the simulation, the possibility of the driver refusing the delivery order is considered. The probability of the driver refusing is represented by a probability distribution and modeled using the beta density function. Finally, the agent uses this probability to perform a Bernoulli test to determine whether to reject the order.

[0012] S3. Agent training: This method combines real-world data and simulated data, and uses the DDQN (Double DeepQ-Learning) algorithm to simultaneously train dual agents, enabling them to make scheduling assignments for the two events of "new order" and "vehicle". In order to solve the overestimation bias problem in traditional Q-learning training, the Double Q-learning algorithm is used.

[0013] Preferably, the S1 specifically comprises the following steps:

[0014] S1-1, indicates the environment:

[0015] The present invention uses a simple coordinate system based on latitude and longitude to represent the location of the vehicle and the destination, representing the environment as a space where delivery orders can be placed;

[0016] S1-2. Proposed event-based SMDP formula:

[0017] SMDPs (Semi-Markov Decision Processes) refers to a set of multiple SMDPs, and the objective function is as follows:

[0018]

[0019] Among them, s k In the decision period t k state, continuous decision period t k and t k+1 The time interval between is a random variable, integrating the reward function over the interval, and e -βt represents the limiting form of the discrete-time discount factor γ for continuous time, where e -β =γ;.

[0020] The present invention defines two important events in the system: "the emergence of new orders" (new order events) and "vehicle capacity vacancy" (vehicle events), which trigger scheduling allocation. In the new order event, the problem is simplified to selecting the most suitable vehicle to deliver the specified new order; in the vehicle event, the problem is simplified to selecting the most suitable series of orders for the specified vehicle to deliver.

[0021] S1-3. Define state and action space:

[0022] The state of the system depends on the triggered event. In the case of a vehicle event, the state is represented by a specific vehicle and the action is determined by the unassigned order; in the case of a new order event, the state is represented by a specific order and the action is established by considering the available vehicles. At the same time, the present invention also adds three context features: resource demand ratio, cycle feature 1, and cycle feature 2. The data structure of the agent is consistent with the specific event, and it can make a wise allocation based on the relevant information.

[0023] S1-4. Define the reward function:

[0024] Agents are rewarded based on the duration of the delivery trip and the time it takes for the vehicle to pick up the goods. The general reward R consists of two parts: When the Agent is responsible for the new order event, Indicates the estimated time required from the origin to the destination. At this time, the hyperparameter b is formulated based on real empirical data and represents the shelf life of fresh products. Low b values ​​encourage the agent to increase the priority of the order, thereby indirectly reducing the total duration of delivery. When the agent is responsible for vehicle events, The hyperparameter b represents the freshness period of fresh products. A low b value encourages the agent to choose a vehicle close to the pickup location, which indirectly reduces the time it takes for the vehicle to pick up goods.

[0025] S1-5. Logical definition in environmental dynamics:

[0026] The environment evolves over time, with state transitions occurring at discrete points in time and remaining in a given state for a random period of time. The environment includes six main sources of randomness: orders appear at random times in the environment, vehicles require different durations to move between locations, vehicles have different availability (including capacity, refrigeration equipment, and other requirements), drivers may reject the dispatch suggestions provided by the agent with a certain probability, customers' maximum tolerable vehicle pickup time has an unknown probability distribution, and different fresh food categories have different shelf life limits.

[0027] Preferably, the classes of the basic entities are defined as Order class and Vehicle class. Order class represents the order request issued by the customer, and its main attributes include creation time, starting point, destination and the maximum time tolerated by the customer for the vehicle to collect goods. Vehicle class represents the vehicle in the simulator, and its main attributes include current location, route planning and vehicle availability (including capacity, refrigeration equipment and other requirements).

[0028] Preferably, the class definition of the Agent is VehicleAgent class and NewOrderAgent class. VehicleAgent class is responsible for selecting orders from the order pool, while NewOrderAgent class is responsible for selecting vehicles. The two Agent classes are very similar, and the main difference is their action sets;

[0029] Preferably, the Environment class is the core of the simulator, which configures multiple simulation parameters (such as the number of vehicles in the fleet, the average speed of the vehicles, and the maximum number of orders per day) and controls the entire fresh food delivery vehicle scheduling process.

[0030] Preferably, when the "new order event" is triggered, NewOrderAgent immediately scans all vehicles in the fleet and selects the most suitable vehicle in terms of availability and capacity. If the selected vehicle is also suitable for the delivery route, the agent will initiate the allocation and start the delivery after the driver confirms. If the selected vehicle is not suitable for the delivery route, the system will place the order in the unresponded order pool until it is assigned to a suitable vehicle.

[0031] Preferably, when the "vehicle event" is triggered, VehicleAgent queries the order pool to determine the order that is suitable for the delivery requirements and cargo size. If the selected order is also suitable for the delivery route, the agent will start the allocation and start picking up and delivering after the driver confirms; if the selected order is not suitable for the delivery route, the system will place the vehicle in the vehicle pool until it is assigned to a suitable order;

[0032] Preferably, the S3 specifically includes the following steps:

[0033] S3-1. Collect data:

[0034] The simulator utilizes the departure place, destination location and arrival time in the real data and uses probability distribution to simulate the data, combining the real world and the simulated data for training the agent of the present invention.

[0035] S3-2, Classification Agent:

[0036] Since in the SMDP formulation of DVDP, allocation occurs in two different types of events, the present invention trains two different agents: NewOrderAgent and VehicleAgent.

[0037] S3-3, sampling conversion:

[0038] At the initial stage of the training process, the agent will randomly take actions and collect a series of experience transitions (represented by tuples: current state, next state, action, reward, event duration) after understanding the consequences of the behavior in the environment. As the simulation progresses, these transitions are stored in a pool called the "Experience Buffer" (EB); in order to ensure the sample diversity required for training, break the temporal correlation and reduce the non-stationarity of the data, the agent randomly selects a batch of experience transitions from the "Experience Buffer" to form a batch.

[0039] S3-4, Deep Neural Network Driver:

[0040] Each agent has two deep neural networks that have the same structure except for the parameters. When the “experience buffer” accumulates a certain number of samples, the experience tuples in the batch are connected to each other, representing the potential allocation between vehicles and orders in a specific context.

[0041] These experience tuples are fed into a gradient step to update the parameters of the deep neural network through the back-propagation algorithm. One of the neural networks performs gradient descent at each step, while the other neural network performs parameter updates after a certain number of steps to control network parameter synchronization.

[0042] Deep neural network combined with DDQN (Double Deep Q-Learning) algorithm, using two functions q A and q B , each q function uses the value of another q function to update the next state to drive the agent to more accurately estimate the Q value of the current state and action, with q A For example:

[0043]

[0044] Where s represents the state of the environment in which the agent is located; a represents the action chosen by the agent in a given state; s' represents the new state that the agent enters after performing action a; r represents the immediate reward obtained by the agent from the environment after performing the action; γ = e -βτ is the discount factor, which indicates the decay rate of the importance of future rewards; α is the learning rate, which indicates the learning rate when updating the Q value, which determines the relative importance of the new estimate to the old estimate when updating the Q value. Indicates that action a with the largest Q value is selected in state s'.

[0045] The agent selects the action with the maximum Q value based on the estimated Q value. NewOrderAgent assigns newly arrived orders to available vehicles, while VehicleAgent schedules vehicles with spare capacity to serve waiting orders.

[0046] (III) Beneficial effects

[0047] Compared with the prior art, the present invention provides a method for optimizing dynamic vehicle scheduling for fresh food delivery based on the DDQN algorithm, which has the following beneficial effects:

[0048] 1. The present invention introduces deep reinforcement learning technology into the field of fresh food delivery vehicle scheduling technology, regards the dynamic vehicle scheduling problem of fresh food delivery as a continuous time process based on the SMDP framework, and greatly shortens the scheduling time of the agent; uses the DDQN algorithm to train the dual agents to make scheduling decisions for the dual events of "new order event" and "vehicle event", which greatly reduces the combinatorial complexity of the decision space, and improves the system resource utilization and scheduling efficiency while considering multiple allocation constraints.

[0049] 2. The present invention improves system resource utilization and scheduling efficiency, thereby ensuring the timeliness of fresh produce and ensuring timely delivery of orders, effectively solving the problem of reduced timeliness of fresh produce due to delayed fresh produce delivery. At the same time, by considering multiple allocation constraints, the diverse demands of orders can be more comprehensively evaluated and met, and efficient allocation can be ensured under the premise of meeting the constraints, effectively solving the problem of damaged fresh produce quality due to unreasonable vehicle allocation such as improper temperature control, insufficient storage space, and unreasonable delivery routes. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is a graphical flow chart of a method for dynamic vehicle scheduling optimization for fresh food delivery based on the DDQN algorithm proposed in the present invention. DETAILED DESCRIPTION

[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0052] Embodiment 1:

[0053] See attached Figure 1 , a dynamic vehicle scheduling optimization method for fresh food delivery based on DDQN algorithm, including the following steps:

[0054] S1. Treat the dynamic vehicle scheduling problem in fresh food delivery as a continuous time process based on the SMDP framework: This method proposes an event-based SMDP formula based on the characteristics that fresh food delivery orders appear randomly over time and the time intervals between continuous allocations are random, and defines the basic components of SMDP: environment, state, action space, reward function and environmental dynamics. In the system, two important events that trigger allocation are clearly defined: "new order event" and "vehicle event", which simplifies the original many-to-many allocation scheduling problem into a one-to-many allocation scheduling problem; specifically, it includes the following contents:

[0055] S1-1, indicates the environment:

[0056] The present invention uses a simple coordinate system based on latitude and longitude to represent the location of the vehicle and the destination, representing the environment as a space where delivery orders can be placed;

[0057] S1-2. Proposed event-based SMDP formula:

[0058] SMDPs (Semi-Markov Decision Processes) refers to a set of multiple SMDPs, and the objective function is as follows:

[0059]

[0060] Among them, s k In the decision period t k state, continuous decision period t k and t k+1 The time interval between is a random variable, integrating the reward function over the interval, and e -βt represents the limiting form of the discrete-time discount factor γ for continuous time, where e -β =γ.

[0061] The present invention defines two important events in the system: "the emergence of new orders" (new order events) and "vehicle capacity vacancy" (vehicle events), which trigger scheduling allocation. In the new order event, the problem is simplified to selecting the most suitable vehicle to deliver the specified new order; in the vehicle event, the problem is simplified to selecting the most suitable series of orders for the specified vehicle to deliver.

[0062] S1-3. Define state and action space:

[0063] The state of the system depends on the triggered event. In the case of a vehicle event, the state is represented by a specific vehicle and the action is determined by the unassigned order; in the case of a new order event, the state is represented by a specific order and the action is established by considering the available vehicles. At the same time, the present invention also adds three context features: resource demand ratio, cycle feature 1, and cycle feature 2. The data structure of the agent is consistent with the specific event, and it can make a wise allocation based on the relevant information.

[0064] S1-4. Define the reward function:

[0065] Agents are rewarded based on the duration of the delivery trip and the time it takes for the vehicle to pick up the goods. The general reward R consists of two parts: When the Agent is responsible for the new order event, Indicates the estimated time required from the origin to the destination. At this time, the hyperparameter b is formulated based on real empirical data and represents the shelf life of fresh products. Low b values ​​encourage the agent to increase the priority of the order, thereby indirectly reducing the total duration of delivery. When the agent is responsible for vehicle events, The hyperparameter b represents the freshness period of fresh products. A low b value encourages the agent to choose a vehicle close to the pickup location, thereby indirectly reducing the vehicle pickup time.

[0066] S1-5. Logical definition in environmental dynamics:

[0067] The environment evolves over time, with state transitions occurring at discrete points in time and remaining in a given state for a random period of time. The environment includes six main sources of randomness: orders appear at random times in the environment, vehicles require different durations to move between locations, vehicles have different availability (including capacity, refrigeration equipment, and other requirements), drivers may reject the dispatch suggestions provided by the agent with a certain probability, customers' maximum tolerable vehicle pickup time has an unknown probability distribution, and different fresh food categories have different shelf life limits.

[0068] S2. Simulation using Discrete Event Simulation (DES): This method uses Python to configure the simulator. The simulator consists of three types of objects: basic entities, agents, and environments. The simulator works by maintaining a list of orders in chronological order and using specific processing routines to handle these events.

[0069] During the simulation, the possibility of the driver refusing the delivery order is considered. The probability of the driver refusing is represented by a probability distribution and modeled using the beta density function. Finally, the agent uses this probability to perform a Bernoulli test to determine whether to reject the order.

[0070] The classes of the basic entities are defined as Order and Vehicle. The Order class represents the order request issued by the customer, and its main attributes include creation time, origin, destination, and the maximum time the customer tolerates for the vehicle to pick up the goods. The Vehicle class represents the vehicle in the simulator, and its main attributes include current location, route planning, and vehicle availability (including capacity, refrigeration equipment, and other requirements).

[0071] The Agent classes are defined as VehicleAgent and NewOrderAgent. VehicleAgent is responsible for selecting orders from the order pool, while NewOrderAgent is responsible for selecting vehicles. The two Agent classes are very similar, and the main difference is their action sets.

[0072] The Environment class is the core of the simulator, which configures multiple simulation parameters (such as the number of vehicles in the fleet, the average speed of the vehicles, and the maximum number of orders per day) and controls the entire fresh food delivery vehicle scheduling process.

[0073] When the "new order event" is triggered, NewOrderAgent immediately scans all vehicles in the fleet and selects the most suitable vehicle in terms of availability and capacity. If the selected vehicle is also suitable for the delivery route, the agent will initiate the allocation and start the delivery after the driver confirms. If the selected vehicle is not suitable for the delivery route, the system will place the order in the unresponded order pool until it is assigned to a suitable vehicle.

[0074] When the "vehicle event" is triggered, VehicleAgent queries the order pool to determine the order that is suitable for the delivery requirements and cargo size. If the selected order is also suitable for the delivery route, the agent will start the allocation and start picking up and delivering after the driver confirms; if the selected order is not suitable for the delivery route, the system will place the vehicle in the vehicle pool until it is assigned to a suitable order.

[0075] S3. Agent training: This method combines real-world data and simulated data, and uses the DDQN algorithm to simultaneously train dual agents, enabling them to make scheduling assignments for the two events of "new order" and "vehicle". In order to solve the overestimation bias problem in traditional Q-learning training, the Double Q-learning algorithm is used;

[0076] The specific contents include the following:

[0077] S3-1. Collect data:

[0078] The simulator utilizes the departure place, destination location and arrival time in the real data and uses probability distribution to simulate the data, combining the real world and the simulated data for training the agent of the present invention.

[0079] S3-2, Classification Agent:

[0080] Since in the SMDP formulation of DVDP, allocation occurs in two different types of events, the present invention trains two different agents: NewOrderAgent and VehicleAgent.

[0081] S3-3, sampling conversion:

[0082] At the initial stage of the training process, the agent will randomly take actions and collect a series of experience transitions (represented by tuples: current state, next state, action, reward, event duration) after understanding the consequences of the behavior in the environment. As the simulation progresses, these transitions are stored in a pool called the "Experience Buffer" (EB); in order to ensure the sample diversity required for training, break the temporal correlation and reduce the non-stationarity of the data, the agent randomly selects a batch of experience transitions from the experience buffer to form a batch.

[0083] S3-4, Deep Neural Network Driver:

[0084] Each agent has two deep neural networks that have the same structure except for the parameters. After the experience buffer accumulates a certain number of samples, the experience tuples in the batch are connected to each other, representing the potential allocation between vehicles and orders in a specific context.

[0085] These experience tuples are fed into a gradient step to update the parameters of the deep neural network through the back-propagation algorithm. One of the neural networks performs gradient descent at each step, while the other neural network performs parameter updates after a certain number of steps to control network parameter synchronization.

[0086] Deep neural network combined with DDQN (Double Deep Q-Learning) algorithm, using two functions q A and q B , each q function uses the value of another q function to update the next state to drive the agent to more accurately estimate the Q value of the current state and action, with q A For example:

[0087]

[0088] Where s represents the state of the environment in which the agent is located; a represents the action chosen by the agent in a given state; s' represents the new state that the agent enters after performing action a; r represents the immediate reward obtained by the agent from the environment after performing the action; γ = e -βτ is the discount factor, which indicates the decay rate of the importance of future rewards; α is the learning rate, which indicates the learning rate when updating the Q value, which determines the relative importance of the new estimate to the old estimate when updating the Q value. Indicates that action a with the largest Q value is selected in state s'.

[0089] The agent selects the action with the maximum Q value based on the estimated Q value. NewOrderAgent assigns newly arrived orders to available vehicles, while VehicleAgent schedules vehicles with spare capacity to serve waiting orders.

[0090] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A distribution vehicle dynamic scheduling optimization method based on DDQN algorithm, characterized in that: The following steps are involved: S1. The dynamic vehicle scheduling problem in fresh food delivery is regarded as a continuous time process based on the SMDP framework: According to the characteristics that fresh food delivery orders appear randomly over time and the time intervals between continuous allocations are random, an event-based SMDP formula is proposed, and the basic components of SMDP are defined: environment, state, action space, reward function and environmental dynamics; in the system, two important events that trigger allocation are clearly defined: "new order event" and "vehicle event", which simplifies the original many-to-many allocation scheduling problem into a one-to-many allocation scheduling problem; S2. Simulate using a discrete event simulator: Use Python to configure a discrete event simulator; use the simulator to maintain a chronological order list, and use specific processing routines to handle "new order events" and "vehicle events": During the simulation, the probability of the driver's refusal is represented by a probability distribution and modeled using the beta density function. Finally, the agent uses this probability to perform a Bernoulli trial to determine whether to reject the order. S3, Agent training: Combine real-world data and simulated data, and use the DDQN algorithm to train the dual agents simultaneously, so that they can make scheduling allocations for "new order events" and "vehicle events", which specifically includes the following steps: S3-1. Collect data: The simulator uses the departure and destination locations and arrival times in real data and uses probability distribution to simulate data, combining real-world and simulated data for agent training. S3-2, Classification Agent: Since the allocation occurs in two different types of events in the SMDP formulation of DVDP, two different agents, NewOrderAgent and VehicleAgent, are trained separately; S3-3, sampling conversion: 1) In the initial stage, the agent learns the consequences of its actions in the environment, makes random actions and collects a series of experience transformations; 2) storing the experience transformation described in step 1) in a pool of "experience buffer"; 3) The agent randomly selects a batch of experience transformations from the "experience buffer" to form a batch to ensure the sample diversity required for training, break the temporal correlation and reduce the non-stationarity of the data; S3-4, Deep Neural Network Driver: When the "experience buffer" accumulates a certain number of samples, the experience tuples in the batch are connected to each other, representing the potential allocation between vehicles and orders in a specific context; The experience tuple is input to perform a gradient step, and the parameters of the deep neural network are updated through a back-propagation algorithm, wherein one neural network performs a gradient descent at each step, and the other neural network performs a parameter update after a certain number of steps to control the synchronization of network parameters; The deep neural network combines the DDQN algorithm and uses two functions q A and q B , each q Function uses another q The value of the function updates the next state to drive the Agent to more accurately estimate the Q value of the current state and action. q A For example: in, s Represents the state of the environment in which the agent is located; a represents the action chosen by the agent in a given state; s 'Indicates that the action is being executed a The new state that the agent enters; r Represents the immediate reward the agent receives from the environment after performing an action; γ = e -βτ is the discount factor, which indicates the decay rate of importance of future rewards; α is the learning rate, which indicates the learning rate when updating the Q value, which determines the relative importance of the new estimate to the old estimate when updating the Q value. Indicates in status s 'Select the action with the largest Q value a ; The agent selects the action with the maximum Q value based on the estimated Q value; NewOrderAgent assigns newly arrived orders to available vehicles; VehicleAgent serves waiting orders with vehicles with spare capacity.

2. According to claim 1, a distribution vehicle dynamic scheduling optimization method based on DDQN algorithm is characterized in that: The S1 specifically includes the following contents: S1-1, indicates the environment: A simple coordinate system based on latitude and longitude is used to represent the location of vehicles and destinations, representing the environment as a space where delivery orders can be placed; S1-2. Proposed event-based SMDP formula: SMDPs refers to a set of multiple SMDPs, and its objective function is as follows: in, s k In the decision-making period t k state, continuous decision period t k and t k+1 The time interval between is a random variable, integrating the reward function over the interval, and e -βt Discrete-time discount factor representing continuous time γ The limiting form of e -β =γ ; Define two important events in the system: "new order event" specifically refers to "new order occurrence", and "vehicle event" specifically refers to "vehicle capacity availability". The new order event and vehicle event trigger scheduling allocation; in the new order event, the problem is simplified to selecting the most suitable vehicle to deliver the specified new order; in the vehicle event, the problem is simplified to selecting the most suitable series of orders for the specified vehicle to deliver; S1-3. Define state and action space: The state of the system depends on the triggered event: in the case of a vehicle event, the state is represented by a specific vehicle and the action is determined by the unassigned order; in the case of a new order event, the state is represented by a specific order and the action is established by considering the available vehicles; three context features are added: resource demand ratio, cycle feature 1, and cycle feature 2; the data structure of the agent is consistent with the specific event, and the allocation is made according to the relevant information; S1-4. Define the reward function: Agents are rewarded based on the duration of the delivery trip and the time it takes for the vehicle to pick up the goods: General rewards R It consists of two parts: , when the Agent is responsible for the new order event, represents the estimated time required to travel from the origin to the destination. b Based on real experience data, it indicates the shelf life of fresh products. b The value encourages the agent to increase the priority of the order, which indirectly reduces the total duration of the delivery; when the agent is responsible for the vehicle event, Represents the time it takes for a vehicle to collect cargo, a hyperparameter b Indicates the shelf life of fresh products, low b The value encourages the agent to choose a vehicle that is close to the pickup location, which indirectly reduces the vehicle pickup time; S1-5. Logical definition in environmental dynamics: Six sources of randomness define the environment: ① Orders appear at random times in the environment; ② The movement of vehicles between locations requires different durations; ③Vehicles have different availability; ④ The driver rejects the dispatch suggestion provided by the Agent with a certain probability; ⑤ The customer’s maximum tolerable vehicle pickup time has an unknown probability distribution; ⑥The shelf life of different fresh products varies.

3. According to claim 1, a distribution vehicle dynamic scheduling optimization method based on DDQN algorithm is characterized in that: The simulator described in S2 consists of three types of objects: basic entities, agents, and environments.

4. According to claim 3, a distribution vehicle dynamic scheduling optimization method based on DDQN algorithm is characterized in that: The classes of the basic entity are defined as Order class and Vehicle class: the Order class represents the order request issued by the customer, and its attributes include creation time, starting point, destination and the customer's maximum tolerance for vehicle pickup time; the Vehicle class represents the vehicle in the simulator, and its attributes include current location, route planning and vehicle availability.

5. The method for dynamic dispatch optimization of distribution vehicles based on DDQN algorithm according to claim 3 is characterized in that: The class definition of the Agent is VehicleAgent class and NewOrderAgent class: the VehicleAgent class is responsible for selecting an order from the order pool; the NewOrderAgent class is responsible for selecting a vehicle.

6. The method for dynamic dispatch optimization of distribution vehicles based on DDQN algorithm according to claim 3 is characterized in that: The environment class is the core of the discrete event simulator, which is used to configure multiple simulation parameters and control the entire fresh food delivery vehicle scheduling process.

7. The method for dynamic dispatch optimization of delivery vehicles based on DDQN algorithm according to claim 5 is characterized in that: When the "new order event" is triggered, NewOrderAgent immediately scans all vehicles in the fleet and selects the vehicle with the most suitable availability and capacity; If the selected vehicle is also suitable for the delivery route, the agent will initiate the allocation and start the delivery after the driver confirms; If the selected vehicle is not suitable for the delivery route, the system will place the order in the unresponse order pool until it is assigned to a suitable vehicle.

8. The method for dynamic dispatch optimization of distribution vehicles based on DDQN algorithm according to claim 5 is characterized in that: When the "vehicle event" is triggered, VehicleAgent queries the order pool to determine the order that is suitable for the delivery requirements and cargo size; if the selected order is also suitable for the delivery route, the intelligent body will start the allocation and start picking up and delivering after the driver confirms; if the selected order is not suitable for the delivery route, the system will place the vehicle in the vehicle pool until it is assigned to a suitable order.

Citation Information

Patent Citations

  • Instant delivery order distribution system for rider-unmanned vehicle cooperative delivery

    CN116415882A

  • Multi-agent reinforcement learning for order-dispatching via order-vehicle distribution matching

    US20200273346A1