A port collection and distribution vehicle dynamic intelligent scheduling method

By constructing a vehicle task selection decision model and combining reinforcement learning and deep learning, the problems of low learning efficiency and the influence of dynamic factors in the dispatching of collection and distribution vehicles in the existing technology are solved, achieving stable and efficient dispatching in dynamic environments and reducing transportation costs and carbon emissions.

CN117132071BActive Publication Date: 2026-05-19DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DALIAN UNIV OF TECH
Filing Date
2023-09-08
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing Q-Learning algorithms have low learning efficiency in the scheduling of collection and distribution vehicles, limiting the improvement of scheduling performance. They cannot effectively cope with dynamic changes in large-scale problems, resulting in scheduling schemes being unsuitable or requiring rescheduling.

Method used

A vehicle task selection decision unit model is constructed using the Res-D3QN algorithm. By combining reinforcement learning and deep learning, the model is divided into two stages: vehicle assignment and dynamic task allocation. The Res-D3QN algorithm is used for action decision-making, and the vehicle scheduling is optimized by combining the experience playback mechanism and residual network.

Benefits of technology

It achieves stability and accuracy in vehicle scheduling under dynamic environments, reduces transportation costs and carbon emissions, and improves the convergence effect and accuracy of scheduling algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117132071B_ABST
    Figure CN117132071B_ABST
Patent Text Reader

Abstract

The application discloses a kind of port collection and distribution vehicle dynamic intelligent scheduling method, comprising: constructing vehicle task selection decision unit model MDP-VTS, the model includes vehicle assignment model and dynamic task allocation model, the relationship between two models is as shown in Figure 2;Wherein vehicle assignment model is in the execution before task start, the vehicle to be executed collection and distribution task is assigned, and dynamic task allocation model is to assign task for vehicle in real time according to the task execution condition of vehicle and overall task information;The action task selection strategy of the vehicle task selection decision unit model MDP-VTS is solved by Res-D3QN algorithm;The solving process is divided into learning stage and application stage, and the training update of Q network parameter is carried out in the learning stage using multiple rounds of incremental learning;The application stage first carries out simulation assignment and optimizes vehicle assignment scheme, and then carries out dynamic task allocation.This method can dynamically formulate adaptive vehicle intelligent scheduling assignment scheme according to task and job environment change.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle scheduling technology, specifically to a dynamic intelligent scheduling method for port collection and distribution vehicles. Background Technology

[0002] With the development of information technology, smart ports have become the mainstream direction for future port upgrades and transformations. A smart port refers to the integration of port-related businesses and management innovations using emerging information technologies such as the Internet, big data, and artificial intelligence. This optimizes port production and operation processes, shifting port development from primarily relying on increased resource input to primarily relying on technological progress and management innovation, and systematically improving port operational efficiency and service value-added. Port cargo handling, as a crucial component of port operations, urgently requires scientific scheduling and management of cargo vehicles, and the rational allocation of vehicle tasks to improve cargo handling efficiency, reduce cargo fleet operating costs, and lower carbon emissions. This is a pressing need for building smart and green ports.

[0003] Q-Learning is an important reinforcement learning algorithm that learns continuously through trial and error with the environment to obtain optimized action selection and decision-making strategies. It has been successfully applied in fields such as automatic control, robotics, and intelligent scheduling. However, Q-Learning algorithms are suitable for solving problems with finite state spaces. For large-scale problems, such as the scheduling of vehicles involved in transportation, this type of algorithm suffers from low learning efficiency and limited improvement in scheduling performance.

[0004] Shipping vehicle dispatching refers to the scheduling and arrangement of container handling vehicles to perform container transportation tasks (including port entry and exit tasks) between freight stations (or consignor / consignee yards) and container terminals, minimizing the total operating cost of the transportation fleet while meeting the requirements of the shipping tasks. Currently, the main technical methods fall into two categories:

[0005] (1) Static scheduling methods based on heuristic algorithms. This type of method establishes a container delivery and pickup scheduling model for collection and distribution vehicles and uses heuristic algorithms to solve the model. For example, a container vehicle scheduling model for a single yard with multiple terminals is constructed with the objectives of minimizing the number of assigned vehicles and minimizing total carbon emissions. An improved ant colony algorithm is designed to solve the model, resulting in a container delivery and pickup scheduling scheme for multiple terminals in the off-port yard divided by time period. With the dual objective of minimizing the total operating time and total operating cost of all collection and distribution vehicles, a vehicle container delivery and pickup optimization model is established, and an improved genetic algorithm is designed to solve the model, effectively reducing the number of vehicles used and the time vehicles spend in port. However, the vehicle and task execution process information on which this type of method is based is usually static and cannot take into account dynamic factors such as delays in vehicle entry and exit from the port and traffic congestion. Moreover, the collection and distribution scheduling operation sequence scheme generated by this type of method may become unsuitable due to dynamic changes such as vehicle delays or traffic congestion, resulting in reduced actual scheduling accuracy, or even the original operation sequence scheme becoming unexecutable, requiring rescheduling.

[0006] (2) Rule-based dynamic scheduling methods. These methods primarily rely on prior knowledge to construct vehicle scheduling rules to allocate vehicles for collection and distribution tasks. They can perform static scheduling or dynamic scheduling based on the dynamic information of vehicle execution. The general process of rule-based dynamic scheduling for collection and distribution vehicles is as follows: initially, each vehicle is assigned a task to be performed according to the scheduling rules; when a vehicle finishes or is about to finish its current task, the scheduling center assigns a new transportation task to the current vehicle based on the dynamic status of the vehicle and task execution, using the scheduling rules, until all collection and distribution tasks are completed. However, these methods require prior knowledge to construct scheduling rules, and typically can only perform local optimization scheduling, significantly limiting the improvement of scheduling accuracy. Summary of the Invention

[0007] The purpose of this invention is to propose a dynamic intelligent scheduling method for port collection and distribution vehicles, which enables the scheduling of collection and distribution vehicles to dynamically formulate adaptive intelligent vehicle scheduling and assignment schemes based on changes in tasks and operating environment; reduce the operating costs of collection and distribution vehicles in completing transportation tasks; and reduce carbon dioxide emissions during the collection and distribution process.

[0008] The basic process of container delivery and pickup operations for port collection and distribution vehicles is as follows: Figure 1 As shown, port cargo handling reservations are confirmed by the terminal, and vehicles should typically arrive at the port within the reserved time slot. The operation of cargo handling vehicles is affected by various dynamic factors, such as real-time road conditions affecting vehicle travel time, terminal and off-port yard operations affecting loading and unloading times, and terminal gate operations affecting gate opening times. As vehicle operations dynamically change, subsequent task assignments need to be dynamically optimized. Dynamic scheduling of cargo handling vehicles is a real-time task assignment process.

[0009] This application proposes a dynamic intelligent scheduling method for port collection and distribution vehicles, including:

[0010] Construct a vehicle task selection decision unit model MDP-VTS, which includes a vehicle assignment model and a dynamic task allocation model. The relationship between the two models is as follows: Figure 2 As shown; the vehicle assignment model selects vehicles to perform the collection and distribution tasks before the task begins, while the dynamic task allocation model allocates tasks to vehicles in real time based on the current task execution status of the vehicles and the overall task information.

[0011] The action task selection strategy of the vehicle task selection decision unit model MDP-VTS is obtained by solving the Res-D3QN algorithm.

[0012] The solution process is divided into a learning phase and an application phase. The learning phase uses multi-round incremental learning to train and update the Q-network parameters. The reinforcement learning part uses the Q-network for action decision-making and continuously acquires training samples to expand the experience pool during the vehicle's container transportation task. The deep learning part uses the samples in the experience pool to continuously update the Res-D3QN network parameters. The application phase first performs simulated assignment to optimize the vehicle assignment scheme, and then performs dynamic task allocation.

[0013] In the preferred approach, the goal of the vehicle assignment model is to minimize the total cost of completing all collection and distribution tasks, as shown in equation (1):

[0014]

[0015] Among them, W(s) t k m t ) indicates that vehicle k is in state s t k Execute task m t The total cost (excluding fixed costs) includes driver fees, fuel costs, carbon emission control costs, and penalty fees, as shown in the following formula:

[0016]

[0017] Wherein: T e (s t k m t ) and T u (s t k m t ) represent vehicle k in state s respectively t k Execute task m tThe time consumption function for no-load and idling driving, with a value in min, can be calculated based on vehicle status and task information; the coefficient v is determined according to the vehicle type, taking v1 for owned vehicles (k∈[1, K1]) and v2 for hired vehicles (k∈[K1+1, K1+K2]); T p (s t k m t ) indicates that vehicle k is in state Execute task m t The time function for deviation from the reservation interval, with values ​​in min; θ is θ1 when the vehicle arrives early and θ2 when it is late;

[0018] In equation (2), m t Based on the state s of vehicle k t k To obtain the formula, see below:

[0019] m t =G(s) t k a t )=G(s t k ,π(s t k (3)

[0020] Among them, G(s) t k a t ) indicates that vehicle k is in state s t k Choose action a t The task to be performed at that time; π(s) t k () is the action selection strategy for the MDP-VTS model;

[0021] The relationship between two adjacent decision moments of vehicle k is that the next moment in the simulation is the moment when the vehicle completes the task corresponding to the action selected in the previous moment. Its state transition is the system state obtained after the vehicle selects the corresponding action in the previous state, i.e., s. t+1 k =F(s) t k a t );

[0022] In the preferred mode, the goal of the dynamic task allocation model is to maximize the cumulative immediate reward, as shown in Equation (4); the dynamic task allocation of a vehicle refers to: when a vehicle completes a port collection or evacuation task and is in an executable task moment (decision moment), an action selection strategy function is determined to allocate a task to be executed for it.

[0023]

[0024] Vehicle k in state s t k Execute task m t The total cost (excluding fixed costs) is:

[0025]

[0026] In the formula, T e′ (s t k m t ), T u′ (s t k m t ) and T p′ (s t k m t ) represent vehicle k in state s respectively t k Execute task m t The time spent idling, driving at idle speed, and deviating from the scheduled range is obtained from feedback on the actual operation of the vehicle; m t Based on the state s of vehicle k t k Obtain, same as formula (3).

[0027] In equation (4), r(s) t k a t ) Calculated according to formula (8), it is related to W′(s) t k m t The relationship is:

[0028] ω1:ω2:ω3=(v+(c F +c R ×ψ)·r e ):(v+(c F +c R ×ψ)·r u ):θ (6)

[0029] The relationship between two adjacent decision moments of vehicle k is that, in actual operation, the next moment is the moment when the vehicle completes the task corresponding to the action selected in the previous moment. Its state transition is the actual system state after the vehicle selects the corresponding action in the previous state, i.e., s. t+1 k =F(s) t k a t ).

[0030] In the preferred embodiment, each state s in the state set S of the vehicle task selection decision unit model includes a 6-dimensional vector description based on factors such as task volume, time, and location, expressed as:

[0031] s=(p1, p2, p3, p4, p5, p6) (7)

[0032] In the formula: p1 is the difference between the port arrival and port departure tasks in the next planned time window; p2 is the sum of the port arrival and port departure tasks in the next planned time window; p3 is the task quantity at the current vehicle location and time (the task quantity within the scheduled time period starting from the current vehicle location); p4 is the proportion of the maximum remaining task quantity among all tasks at the current time to the total remaining task quantity; p5 is the total number of vehicles ending at the current location at the current time; p6 is the total number of tasks exceeding the scheduled time at the current time; the variables of each dimension of the state vector are processed using the Z-Score normalization method.

[0033] In the preferred approach, the vehicle action group of the vehicle task selection decision unit model is designed based on the rules of vehicle task screening strategy combination. The main vehicle task screening strategies are: 1) screening tasks with the most urgent reservation time; 2) screening tasks whose starting point is closest to the vehicle's location; 3) screening tasks with the shortest travel time from the starting point to the destination; and 4) screening tasks with the most remaining cargo space. These four task screening strategies are combined and unreasonable combinations are removed to design the task selection action group for vehicle scheduling, consisting of six specific actions as follows:

[0034] 1) a1 - Select the task with the most urgent appointment time that is closest to the vehicle's location from the starting point;

[0035] 2) a2 - Select the tasks with the most urgent appointment times and the shortest travel time from the starting point to the destination;

[0036] 3) a3 - Select the task with the most urgent appointment time and the largest remaining box quantity;

[0037] 4) a4 - Select the task with the shortest time commitment and the closest starting point to the vehicle's location;

[0038] 5) a5 - Select the task with the shortest travel time from the starting point to the destination, based on the task set where the starting point is closest to the vehicle's location;

[0039] 6) a6 - Select the task with the most remaining boxes in the task set closest to the vehicle's location from the starting point.

[0040] In the preferred embodiment, the immediate reward function of the vehicle task selection decision unit model consists of two parts: task reward and time reward, and its formula is as follows:

[0041] r = rd +r t (8)

[0042] r d =λ1r1+λ2r2 (9)

[0043] r t =-(ω1T) e +ω2T u +ω3T c (10)

[0044] In equation (8), r represents immediate return, r d For the task reward, r t For time-based feedback; where r1 is the task urgency feedback value, positive feedback is given if the task reduces the urgency, otherwise negative feedback is given; r2 is the task balance feedback value, positive feedback is given if the task reduces the difference in the remaining port collection and distribution tasks, otherwise negative feedback is given; λ1 and λ2 are both task reward component coefficients; T e T u T c These represent the vehicle's idle time, idling time, and deviation from the scheduled time period during the mission, all in minutes. ω1, ω2, and ω3 are time-reward component coefficients, all in minutes. -1 The emphasis on each item's return can be adjusted by changing the values ​​of the coefficients for each item.

[0045] In the preferred mode, the vehicle k is determined based on the actual operation of the collection and distribution vehicles or the simulated environment. Next, execute action a t The subsequent state can be represented as:

[0046] In the preferred embodiment, the action selection strategy π of the vehicle task selection decision unit model is determined by a strategy function based on the current state of vehicle k. The next action to be selected is:

[0047]

[0048]

[0049] Where r t a Indicates the vehicle is in a certain state. The immediate reward value for executing action 'a'; Indicates vehicle status The cumulative reward for executing action 'a'; γ is the cumulative reward discount factor. The action policy π will be learned using the Res-D3QN algorithm.

[0050] In the preferred approach, the agent's trial-and-error learning method in the Res-D3QN algorithm involves interacting with the environment: Given the current state vector s of a vehicle in the transportation system, the agent uses a Q-Net network with parameters w to evaluate the cumulative reward value Q(s, a; w) for each action performed in that state. It then learns and explores strategies to output action a. After the transportation system executes the agent's action a, the system state transitions from s to s′, and the immediate reward is r. The experience gained during the generation of the immediate reward r is defined as e = (s, a, r, s′), and stored in the experience pool for subsequent training. Simultaneously, it receives immediate reward signals from the environment to initiate a new cycle, such as... Figure 3 As shown.

[0051] The Res-D3QN algorithm uses a deep Q-network to fit the state-action value function. The predicted Q-value is represented as Q(s, a; w, α, β). The value function Q can be further divided into two parts: the first part is the state value function, which depends only on the vehicle state and not on the actions performed by the vehicle, denoted as V(s; w, α); the second part is the advantage function, which depends on both the state and the action, denoted as... These two parts constitute the penultimate layer of the Q-network, and the value function of the output layer of the last layer of the Q-network is the sum of these two parts:

[0052]

[0053] Where w is the network parameter of the common part, α is the network parameter unique to the state value function part, and β is the network parameter unique to the advantage function part.

[0054] The Res-D3QN algorithm decouples action selection and evaluation, using a Q-network to select actions and a target Q-network to determine the value of the actions. The target value that the Q-network needs to approximate when iteratively updating its parameters is y = r + γQ(s′, argmaxQ(s′, a′; w, α, β); w′, α′, β′), where γ is a discount factor. This target value is predicted by another target network with the same structure as the Q-network. The predicted value and the target value are substituted into the loss function to calculate the loss value, as shown in Equation (14).

[0055] Loss(w) = E (a,r,s,r′) [(yQ(s,a;w,α,β)) 2 (14)

[0056] The above loss function, when its partial derivative is taken with respect to the weights w, achieves backpropagation. Gradient descent is used to update the Q-network parameters. During training, only the Q-network parameters are updated initially. After a certain interval, the Q-network parameters are copied to the target network. This breaks the correlation between the Q-network and the target network, improving training stability. The model's network structure is as follows: Figure 4 As shown.

[0057] In the preferred approach, considering the characteristics of the vehicle dynamic scheduling problem, a residual MLP network with linear layers as weight layers was designed as the Q-network based on the residual mechanism. The network structure is as follows: Figure 5 As shown, the network mainly includes the following layers: Linear Layer: All nodes at both ends of the linear layer are fully connected, and the connections between nodes are assigned weights, which are continuously adjusted and updated during training; Nonlinear Mapping Layer: The main function of the nonlinear mapping layer is to perform nonlinear mapping on the output of the linear layer using a nonlinear activation function, employing the ReLU function, i.e., f(x) = max(0, x); Batch Normalization Layer: Batch normalization standardizes the data, facilitating neural network computation; Dropout Layer: The Dropout layer is used to reduce overfitting in the neural network; Residual Block: The residual mechanism alleviates the gradient vanishing problem caused by increasing the depth of the neural network. The residual block structure is shown in the figure. Figure 6 As shown;

[0058] The output of the residual block is: H(x) = F(x) + x, where x is the input of the residual block; H(x) is the output of the residual block; and F(x) is the difference between the output and the input, i.e., the residual. The residual network transforms the identity mapping to be fitted into the fitted residual, that is, when F(x) is 0, the identity mapping H(x) = x is formed, and the training objective of the neural network becomes to approximate the residual F(x) to 0.

[0059] For the state set, action set, and immediate reward of the vehicle task selection decision unit model, the input of the Q network is designed as six dimensions of the state, namely p1 to p6. The output of the Q network is designed as the Q value of the six actions corresponding to the state, calculated according to equation (13). The Q value is used for action decision-making, and the immediate reward is used to quantify the quality of action execution. The input and output design of the Q network for the dynamic scheduling of gathering and dispersing is as follows: Figure 7 As shown.

[0060] Traditional Q-learning action exploration strategies typically employ an ε-greedy exploration strategy. During the learning process, the agent randomly selects actions with probability ε (ε∈[0,1]) and chooses the optimal action (the action that maximizes the Q-value) with probability (1-ε). However, as the learning process progresses, the higher probability of random action selection in the later stages hinders convergence. To address this issue, this invention employs a method where ε and α decrease continuously with the number of learning iterations to ensure convergence. In the early stages of learning, the agent tends to randomly select actions to fully explore the state space, and a larger proportion of new attempts are accepted, allowing for rapid iteration and updating of the Q-value. As the number of learning iterations increases, the probability of the agent selecting the optimal action continuously increases, and a larger proportion of the fully learned Q-value is retained, ensuring both the convergence effect and convergence speed of the algorithm. The improved ε-greedy exploration strategy is as follows:

[0061]

[0062]

[0063] Where τ is the number of learning iterations; ε0 is the initial value of ε; and ζ is the decay coefficient.

[0064] In the preferred approach, the Q-network parameters are trained and updated during the learning phase. The reinforcement learning part utilizes the Q-network for action decisions, continuously acquiring training samples to expand the experience pool as the vehicle performs container transport tasks. The deep learning part then continuously updates the network parameters using samples from the experience pool. These two processes operate in parallel, as follows: Figure 8 As shown, this can be achieved using a multi-round incremental learning approach. The steps are as follows:

[0065] Step 1: Initialize the Q-Net and Target Q-Net network parameters;

[0066] Step 2: Initialize environment parameters: information such as vehicles, task sequences, docks, and freight stations;

[0067] Step 3: Idle vehicles select an action based on their current status and execute the corresponding container transport task;

[0068] Step 4: Calculate the immediate reward for this task and store the experience information from this decision into the experience pool;

[0069] Step 5: Determine if all tasks in the task sequence have been completed. If yes, proceed to Step 2; otherwise, proceed to Step 3.

[0070] Step 6: Randomly select a batch of training samples from the experience pool to train Q-Net and update its network parameters;

[0071] Step 7: Determine if the time point for updating the Target Q-Net network parameters has been reached. If yes, proceed to Step 8; otherwise, proceed to Step 9.

[0072] Step 8: Copy the network parameters of Q-Net to the Target Q-Net;

[0073] Step 9: Determine if the termination criterion is met, i.e., the objective function value converges and the Q-Net network parameters tend to stabilize; if yes, output Q-Net and end the learning phase; otherwise, go to Step 6.

[0074] In the preferred approach, the application phase process steps are as follows:

[0075] Step 1: Simulate vehicle assignment and optimize the vehicle assignment scheme. The process is as follows: Figure 9 As shown;

[0076] Step 1-1: The scheduling center loads the Q-Net network after learning, which was output during the learning phase;

[0077] Step 1-2: Initialize environment parameters and select the vehicle assignment scheme to be tried, i.e. the number of owned and hired vehicles to be dispatched.

[0078] Steps 1-3: Obtain the status of idle vehicles, and the dispatch center assigns the best task to the vehicle based on the learned Q-Net;

[0079] Steps 1-4: The vehicle simulates the execution of this task and sends the detailed work information to the dispatch center; determine whether all tasks have been completed. If yes, proceed to Step 1-5; otherwise, proceed to Step 1-3.

[0080] Step 1-5: Record the total cost of the vehicle assignment scheme, and determine whether all vehicle assignment schemes to be tried have been simulated and calculated. If yes, output the vehicle assignment scheme corresponding to the minimum total cost and end; otherwise, go to Step 1-2.

[0081] Step 2: Dynamic task allocation, the process is as follows: Figure 10 As shown;

[0082] Step 2-1: The scheduling center loads the learned Q-Net output from the learning phase;

[0083] Step 2-2: Initialize environmental parameters and start port collection and distribution operations using the vehicles selected in Step 1;

[0084] Steps 2-3: The vehicle sends a task request to the dispatch center; the dispatch center assigns the best task to the vehicle based on its status.

[0085] Steps 2-4: The vehicle completes the task and sends detailed operation information (operation time, fuel consumption, etc.) to the dispatch center; the dispatch center evaluates the task, calculates the immediate feedback, and adaptively updates the Q-Net network parameters;

[0086] Step 2-5: Determine if all tasks have been completed. If yes, output the total cost and end the scheduling; otherwise, go to Step 2-3.

[0087] The advantages of the above technical solutions adopted in this invention compared with the prior art are as follows:

[0088] (1) The collection and distribution operations between the yard and the wharf of the port's freight station are characterized by long transportation distances, long task cycles, and large workloads. Therefore, vehicles are easily affected by dynamic factors in the environment during operation, causing fluctuations in operation time and affecting subsequent scheduling. This invention divides the dynamic scheduling of collection and distribution vehicles into two stages: simulated vehicle assignment and dynamic task allocation, and establishes a two-stage model. It considers the influence of various dynamic factors to perform real-time task allocation for vehicles, which can provide decision support for the collection and distribution fleet to formulate scheduling plans and vehicle scheduling assignments, and ensure the stability of the scheduling process.

[0089] (2) The large number of dimensions of state variables and the large size of the state space result in an excessively large Q-value table that stores a large number of unupdated Q-values. Furthermore, it cannot be guaranteed that all states are fully learned, and states that are not fully learned or explored can easily lead to significant decision-making biases. Therefore, this invention proposes a method for action decision-making of an agent by using reinforcement learning supplemented by a deep neural network to fit the value function. The deep neural network is used to replace the Q-value table to predict the Q-values, thereby guiding the agent's action decision-making.

[0090] (3) Deep neural networks may become unstable or even fail to converge when fitting the state-action value function. To address this issue, this invention employs an experience replay mechanism and a fixed target network to stabilize the learning effect. The experience replay mechanism improves sample utilization and breaks the correlation between consecutive samples by repeatedly sampling historical samples; the fixed target network can predict state values ​​independently, overcoming the limitation of the Q-network in the ordinary DQN algorithm that only considers the influence of actions, and is therefore applicable to different scenarios.

[0091] (3) The DQN algorithm directly outputs the state-action value Q(s, a; w) of each action under vehicle state s, ignoring the value prediction of state s itself, which causes a certain deviation in the result. To address this problem, this invention improves the output layer of the Q network in the DQN algorithm based on the Dueling DQN algorithm, dividing the value function Q into two parts: the state value function (related to the vehicle state but not to the actions performed by the vehicle) and the advantage function (related to both state and action). In addition, to increase the recognizability of these two parts, the calculation method of the value function of the output layer is changed.

[0092] (4) The present invention adopts an experience replay mechanism, sets up a replay experience pool to store the experience obtained by the vehicle at each decision moment as a sample source for neural network training. When training the neural network, a batch of experience data is randomly selected from the replay experience pool as training samples, and the network parameters are updated by using a gradient descent mechanism. By repeatedly and randomly sampling historical samples, the sample utilization rate is improved and the correlation between continuous samples is broken.

[0093] (5) The target Q-value in DQN is obtained through maximization. Although this method can quickly bring the Q-value closer to the target, it can easily cause the predicted value of the value function to be larger than the true value, resulting in overestimation and affecting vehicle action decisions. This invention decouples action selection and measurement based on the improved idea of ​​Double DQN. It uses a Q-network to select actions and a target Q-network to determine the value of actions. Through this feedback mechanism, the Q-value can be estimated better.

[0094] (6) The gradient vanishing caused by increasing the depth of the neural network results in non-convergence or slow convergence. To address this problem, this invention introduces a residual mechanism to construct a residual MLP network as a Q-network. The residual blocks inside the network use skip connections to alleviate the gradient vanishing problem caused by increasing the depth of the neural network.

[0095] (7) The action exploration strategy of the DQN algorithm generally selects the action that maximizes the Q value as the optimal action. However, as the learning process progresses, the probability of randomly selecting actions is relatively high in the later stages, which is not conducive to convergence. To address this issue, this invention adopts a method of continuously decreasing the action selection probability as the number of learning iterations increases to ensure convergence performance. As the number of learning iterations increases, the probability of the agent selecting the optimal action continuously increases, and the network parameter updates tend to stabilize after sufficient learning, thus ensuring the convergence performance and convergence speed of the algorithm. Attached Figure Description

[0096] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0097] Figure 1 A basic process diagram of port loading and unloading operations;

[0098] Figure 2 A diagram showing the relationship between the vehicle assignment model and the dynamic task allocation model;

[0099] Figure 3 This is a basic framework diagram of the Res-D3QN algorithm;

[0100] Figure 4 This is a diagram of the neural network structure.

[0101] Figure 5 This is a diagram of the residual Q-network structure;

[0102] Figure 6 This is a diagram of the residual block structure.

[0103] Figure 7 Design diagram of Q-Net input and output;

[0104] Figure 8 A flowchart for the learning phase;

[0105] Figure 9 A flowchart simulating vehicle assignment;

[0106] Figure 10 Flowchart for dynamic task assignment. Specific implementation methods

[0107] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit the application; that is, the described embodiments are only a part of the embodiments of this application, and not all of them.

[0108] R-Deep Q Network (DQN) is a relatively effective algorithm combining Q-learning and deep learning. This invention addresses the dynamic scheduling problem of port transport vehicles by introducing a residual network mechanism based on the DQN algorithm. Combining Dueling and Double DQN algorithms, a novel Res-D3QN algorithm is proposed, achieving dynamic intelligent scheduling of port transport vehicles and effectively improving the accuracy and convergence of the scheduling algorithm.

[0109] Considering the dynamic factors of port entry and exit reservations and operations, this invention decomposes the dynamic scheduling problem of entry and exit vehicles into two sub-problems: vehicle assignment and dynamic task allocation. Its main features are: (1) Entry and exit vehicles are shared among multiple terminals, multiple yards, and multiple tasks. Vehicles are not limited to fixed routes and can continuously undertake port entry and exit tasks; (2) Entry and exit vehicles are shared, and the vehicles that can be assigned include vehicles owned by the fleet and vehicles hired from outside. When the fleet's own vehicles cannot meet the task requirements, vehicles from other fleets can be hired to assist in the execution of the task; (3) Dynamic scheduling of vehicle tasks is adopted. After a vehicle completes a task, a new task is assigned in real time based on the dynamic information of the vehicle and the task.

[0110] The conditions for constructing the vehicle task selection decision unit model are: (1) the information of the collection and distribution tasks under the reservation mode is known (including self-owned tasks and external service tasks), specifically including: the starting point, ending point, reservation start and end time, and task container quantity (40TEU containers or two 20TEU containers); (2) the arrival time of the vehicle should be within the reservation period, that is, the arrival time of the vehicle performing the collection task or the arrival time of the vehicle performing the distribution task is constrained by the reservation period; otherwise, additional penalties will be imposed; (3) the vehicle operation time includes: the driving time on the collection and distribution road, the operation time at the terminal, the operation time in the yard outside the port, and the queuing time at the terminal gate; (4) the idling scenarios of the vehicle include: the operation of the vehicle at the terminal, the operation in the yard outside the port, and queuing; (5) the total operation cost of the collection and distribution fleet includes fuel consumption cost, driver cost, penalty cost for exceeding the reservation period, and carbon emission control cost. The above costs are directly proportional to the vehicle's operation time. Vehicle fuel consumption is divided into heavy load fuel consumption, no load fuel consumption, and idling fuel consumption. (6) Vehicles are divided into owned vehicles and hired vehicles. When there are not enough owned vehicles, hired vehicles can be used. There is no limit to the number of hired vehicles. The costs are similar, but the cost of hired vehicles needs to be considered for excess costs. (7) Each vehicle can only undertake one collection and distribution task at a time.

[0111] The proposed dynamic scheduling method for collection and distribution vehicles mainly includes: a vehicle task selection decision unit model (MDP-VTS) and an improved DQN (Deep Q-Network Reinforcement Learning) algorithm (Res-D3QN) for vehicle task selection strategy. Key technical points of the MDP-VTS model include: state vector, action group, immediate reward function, state transition function, and action task selection strategy; the action task selection strategy is learned and solved using the Res-D3QN algorithm. Key technical points of Res-D3QN include: network structure, exploratory learning strategy, and basic learning algorithm flow. The overall process of solving the dynamic scheduling problem for collection and distribution vehicles using the above method is divided into a learning phase and an application phase: In the learning phase, the Res-D3QN network parameters for the vehicle task selection strategy are trained and updated using actual or simulated samples; in the application phase, the learned vehicle task selection strategy (Res-D3QN network) first optimizes the vehicle assignment scheme based on the collection and distribution task plan information, and then dynamically allocates the collection and distribution tasks to be performed by the vehicles based on the real-time state information of the vehicles during actual task execution.

[0112] Based on the scheduling data between a container terminal and a freight station in a certain port, an experiment was conducted using the above method, and its beneficial effects were analyzed.

[0113] A certain cargo handling fleet owns 30 vehicles. Its 20 port handling tasks are generated as follows: the task cycle is 24 hours (0-24 hours), and the starting and ending points of the port handling tasks are randomly selected from the off-port storage yard. Task reservation time slots are randomly selected from the reservation cycle, ranging from 6 to 12 hours. The container volume within each reservation time slot is allocated proportionally to the time slot length. The total task volume of the port handling task sequence is randomly generated from the interval [E-5, E+5]. The E values ​​of the 20 examples are randomly selected from the set {750, 800, 850, 900, 950, 1000}. Ten examples are used to train the Q-network during the learning phase, numbered (T-1) to (T-10); the other ten examples are used to verify the transfer application effect of the trained Q-network during the application phase, numbered (V-1) to (V-10).

[0114] The vehicle's fuel consumption under heavy load is 0.36 L / km, under no-load fuel consumption is 0.24 L / km, and idling fuel consumption is 0.05 L / min. The fuel price is 7.6 yuan / L, the carbon emission coefficient is 1.6 kg / L, the unit cost of carbon emission treatment is 0.25 yuan / kg, the hourly wage for owned vehicle drivers is 15 yuan / h, and the hourly wage for hired vehicle drivers is 20 yuan / h. The average service time at the terminal gate is 2 minutes, the average driving time within the terminal is 3 minutes, and the average service time for the terminal cranes is 3 minutes. The average driving time for vehicles in the off-port storage yard is 3 minutes, and the average service time for the off-port storage yard cranes is 3 minutes. The discount factor γ is 0.85, and the α0 in the exploration strategy... The values ​​are 0.35 and 0.004, respectively.

[0115] To verify the adaptability of the proposed method to dynamic environments, this embodiment considers the interference of dynamic factors on four types of operation times: vehicle travel time, terminal gate service time, terminal container loading and unloading time, and off-port container loading and unloading time, as well as the additional maintenance time required due to vehicle breakdowns during transportation, as shown in Table 1. The operation time after considering dynamic factors is calculated by adding the original operation time to the dynamic time length. Where μ1 and μ2 are the upper and lower boundaries of the dynamic time length, respectively, in minutes; This represents the probability value of the task time changing dynamically.

[0116] Table 1 List of dynamic factors

[0117]

[0118] The Res-D3QN algorithm is compared with the ordinary DQN algorithm and a rule-based scheduling algorithm (i.e., assigning transportation tasks to vehicles using specific rules during the collection and distribution process, with six scheduling rules representing six actions). The results are as follows:

[0119] Table 2 Two-Stage Scheduling Results - Cost Table

[0120]

[0121] Table 3. Two-stage scheduling results - Carbon emission table

[0122]

[0123] Table 4. Results of the DQN algorithm and its comparison with the Res-D3QN algorithm.

[0124]

[0125]

[0126] Table 5. Rule-based scheduling results (variable cost / yuan, variable carbon emissions / kg)

[0127]

[0128] Table 6. Comparison of the optimization level of the Res-D3QN algorithm with rule-based scheduling algorithms.

[0129]

[0130] The comparative results show that the Res-D3QN algorithm outperforms the ordinary DQN algorithm and rule-based scheduling algorithms in computational examples with different task sizes. Compared with the ordinary DQN algorithm, the variable cost and variable carbon emissions are reduced by 10.79% and 11.93% respectively, indicating that the improved strategy for the DQN algorithm proposed in this invention achieves a significant performance improvement; compared with the rule-based scheduling algorithm, the variable cost and variable carbon emissions are reduced by 52.72% and 74.90% respectively.

[0131] In this application, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0132] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A dynamic intelligent scheduling method for port collection and distribution vehicles, characterized in that, include: Construct a vehicle task selection decision unit model, which includes a vehicle assignment model and a dynamic task allocation model; The vehicle assignment model selects vehicles to perform collection and distribution tasks before the task begins, while the dynamic task allocation model allocates tasks to vehicles in real time based on the current task execution status and overall task information. The action task selection strategy of the vehicle task selection decision unit model is obtained by solving the Res-D3QN algorithm. The solution process is divided into a learning phase and an application phase. The learning phase uses multi-round incremental learning to train and update the Q-network parameters. The reinforcement learning part uses the Q-network for action decision-making and continuously acquires training samples to expand the experience pool during the vehicle's container transportation task. The deep learning part uses the samples in the experience pool to continuously update the Res-D3QN network parameters. The application phase first performs simulated assignment to optimize the vehicle assignment scheme, and then performs dynamic task allocation. The action selection strategy of the vehicle task selection decision unit model The vehicle's state is determined by the policy function. In state The next action to be selected is: (11) (12) in Indicates the vehicle is in a certain state. Next action The immediate return value; Indicates vehicle status Next action Cumulative returns; This is a cumulative return discount factor; this action strategy It will be learned through the Res-D3QN algorithm; In the Res-D3QN algorithm, the agent learns through trial and error by interacting with the environment: the agent learns from the current state vector of the vehicles in the transportation system. Using parameters The Q-Net network evaluates the cumulative reward value of performing each action in this state. Q ( s , a ; w ), through learning and exploring strategies to output actions. a The collection and distribution system executes the actions of intelligent agents. a Afterwards, the system status changed from Transfer to The system immediately reported as Generates immediate returns r The experience gained in the process is defined as It is stored in the experience pool for subsequent training; at the same time, it receives immediate feedback signals from the environment to start a new cycle. The value function Q is divided into two parts: the first part is the state value function, which depends only on the vehicle state and not on the actions performed by the vehicle, denoted as... The second part is the advantage function, which is related to both the state and the action, denoted as... These two parts constitute the penultimate layer of the Q-network, and the value function of the output layer of the last layer of the Q-network is the sum of these two parts: (13) in, w These are network parameters for the common parts. These are network parameters unique to the state-value function. These are network parameters unique to the advantage function part; The Res-D3QN algorithm decouples action selection and evaluation, using a Q-network to select actions and a target Q-network to determine the value of each action. The Q-network iteratively updates its parameters to approximate a target value. , As a discount factor, this target value is predicted by another target network with the exact same structure as the Q network; the predicted value and the target value are substituted into the loss function to calculate the loss value, as shown in Equation (14): (14) The above loss function applies to the weights w Finding the partial derivative is equivalent to backpropagation. The gradient descent mechanism is used to update the parameters of the Q network. During training, only the parameters of the Q network are updated first, and after a certain interval, the parameters of the Q network are copied to the target network. A residual MLP network with linear layers as weights was designed as the Q-network based on the residual mechanism. The output of its residual part is as follows: ,in, x For the input of the residual block; H( x ) is the output of the residual block; F( x The difference between the output and the input is called the residual.

2. The method for dynamic intelligent scheduling of port collection and distribution vehicles according to claim 1, characterized in that, The goal of the vehicle assignment model is to minimize the total cost of completing all collection and distribution tasks, as shown in equation (1): (1) in, Indicates vehicle k In state Execute the task The total cost, including driver fees, fuel costs, carbon emission control costs, and penalty fees, is calculated using the following formula: (2) in: and Representing vehicles k In state Execute the task The time consumption function for no-load and idling driving, with values ​​in minutes; coefficients Values ​​are assigned based on vehicle type; for owned vehicles... ,Pick When hiring vehicles ,Pick ; Indicates vehicle In state Execute the task The time function representing the deviation from the reservation interval, with values ​​in min; Pick up when the vehicle arrives at the port ahead of schedule ; take it when late ; In formula (2) According to the vehicle k status To obtain the formula, see below: (3) in, Indicates vehicle k In state Select action The tasks to be performed at that time; Choose the action selection strategy for the decision unit model for vehicle tasks.

3. The method for dynamic intelligent scheduling of port collection and distribution vehicles according to claim 1, characterized in that, The goal of the dynamic task allocation model is to maximize the cumulative immediate reward, as shown in equation (4); (4) vehicle k In state Execute the task The total cost is: (5) In the formula, , and Representing vehicles k In state Execute the task The time spent idling, driving at idle speed, and deviating from the scheduled range is obtained from feedback on the actual operation of the vehicle. According to the vehicle k status get; Immediate feedback and The relationship is: (6)。 4. The method for dynamic intelligent scheduling of port collection and distribution vehicles according to claim 1, characterized in that, The state set of the vehicle task selection decision unit model Each state The expression is: (7) In the formula: This represents the difference between the volume of port arrivals and port departures within the next planned time window. This is the sum of the port arrival and port departure tasks within the next planned time window; The task load at the vehicle's current location and current time; This represents the proportion of the maximum remaining task quantity among all tasks at the current time to the total remaining tasks. This represents the total number of vehicles at the current time whose destination is the current location. This represents the total number of tasks that have exceeded the scheduled time. The vehicle action group of the vehicle task selection decision unit model is designed based on the rule combination of vehicle task screening strategies. The six specific actions are as follows: 1) —Select the tasks with the most urgent appointment times and the ones closest to the vehicle's location from the starting point; 2) —Select the tasks with the most urgent appointment times and the shortest travel time from the start to the finish line; 3) —Select the task with the most urgent appointment time and the largest remaining box quantity. 4) —Select the task with the shortest time commitment and the closest task to the vehicle's location from the starting point; 5) —Select the task with the shortest travel time from the starting point to the destination, based on the task set whose starting point is closest to the vehicle's location; 6) —Select the task with the most remaining boxes in the task set closest to the vehicle's location.

5. The method for dynamic intelligent scheduling of port collection and distribution vehicles according to claim 3, characterized in that, The immediate reward function of the vehicle task selection decision unit model consists of two parts: task reward and time reward. Its formula is as follows: (8) (9) (10) In equation (8), For immediate return, In return for the task, For the reward of time; among which, This is a feedback value for task urgency. If the urgency of the task is reduced, positive feedback is given; otherwise, negative feedback is given. This is the task balance feedback value. If the current task reduces the difference in the remaining port collection and distribution task volume, positive feedback is given; otherwise, negative feedback is given. , All are task reward component coefficients; , , These represent the vehicle's idle time, idling time, and the time of the deviating from the scheduled time during the mission, all in minutes. , , These are time-based return factor values, all in minutes. -1 .

6. The method for dynamic intelligent scheduling of port collection and distribution vehicles according to claim 1, characterized in that, The exploration strategy is as follows: (15) (16) in, Number of times; for The initial value; This is the attenuation coefficient.

7. The method for dynamic intelligent scheduling of port collection and distribution vehicles according to claim 1, characterized in that, The learning phase steps are as follows: Step 1: Initialize the Q-Net and Target Q-Net network parameters; Step 2: Initialize environment parameters: vehicle, task sequence, dock and freight station information; Step 3: Idle vehicles select an action based on their current status and execute the corresponding container transport task; Step 4: Calculate the immediate reward for this task and store the experience information from this decision into the experience pool; Step 5: Determine if all tasks in the task sequence have been completed. If yes, proceed to Step 2; otherwise, proceed to Step 3. Step 6: Randomly select a batch of training samples from the experience pool to train Q-Net and update its network parameters; Step 7: Determine if the time point for updating the Target Q-Net network parameters has been reached. If yes, proceed to Step 8; otherwise, proceed to Step 9. Step 8: Copy the network parameters of Q-Net to the Target Q-Net; Step 9: Determine if the termination criterion is met, i.e., the objective function value converges and the Q-Net network parameters tend to stabilize; if yes, output Q-Net and end the learning phase; otherwise, go to Step 6.

8. The method for dynamic intelligent scheduling of port collection and distribution vehicles according to claim 7, characterized in that, The application phase steps are as follows: Step 1: Simulate vehicle assignment and optimize the vehicle assignment scheme; Step 1-1: The scheduling center loads the Q-Net network after learning, which was output during the learning phase; Step 1-2: Initialize environment parameters and select the vehicle assignment scheme to be tried, i.e. the number of owned and hired vehicles to be dispatched. Steps 1-3: Obtain the status of idle vehicles, and the dispatch center assigns the best task to the vehicle based on the learned Q-Net; Steps 1-4: The vehicle simulates the execution of this task and sends the detailed work information to the dispatch center; determine whether all tasks have been completed. If yes, proceed to Step 1-5; otherwise, proceed to Step 1-3. Step 1-5: Record the total cost of the vehicle assignment scheme, and determine whether all vehicle assignment schemes to be tried have been simulated and calculated. If yes, output the vehicle assignment scheme corresponding to the minimum total cost and end; otherwise, go to Step 1-2. Step 2: Dynamic task allocation; Step 2-1: The scheduling center loads the learned Q-Net output from the learning phase; Step 2-2: Initialize environmental parameters and start port collection and distribution operations using the vehicles selected in Step 1; Steps 2-3: The vehicle sends a task request to the dispatch center; the dispatch center assigns the best task to the vehicle based on its status. Steps 2-4: The vehicle completes the task and sends the detailed operation information to the dispatch center; the dispatch center evaluates the task, calculates the immediate feedback, and adaptively updates the Q-Net network parameters; Step 2-5: Determine if all tasks have been completed. If yes, output the total cost and end the scheduling; otherwise, go to Step 2-3.