Container connection transportation path planning method based on deep reinforcement learning

By constructing a mixed-integer programming model and using the Transformer algorithm to optimize fuel consumption, a container relay transportation route planning method based on deep reinforcement learning is proposed. This solves the problem of insufficient decision-making in complex scenarios in existing technologies and achieves efficient container relay transportation.

CN122048209APending Publication Date: 2026-05-15ZHONGYUAN ENGINEERING COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHONGYUAN ENGINEERING COLLEGE
Filing Date
2026-01-26
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing optimization methods for container intermodal transport are unable to effectively describe complex spatial topological relationships and inherent symmetries, resulting in insufficient decision quality and generalization ability of intelligent optimization algorithms in high-dimensional scenarios, which fails to effectively reduce transportation costs and improve efficiency.

Method used

A mixed-integer nonlinear programming model is constructed and linearized using a deep reinforcement learning approach. Combined with the Transformer deep reinforcement learning algorithm, multi-head attention is used to capture the dependencies of the container docking transportation network and optimize route planning to minimize fuel consumption.

Benefits of technology

In large-scale and complex scenarios, it improves decision-making quality and generalization ability, reduces transportation carbon emissions, adapts to mixed-load scenarios of multi-size containers, and improves the efficiency of container connection and transportation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122048209A_ABST
    Figure CN122048209A_ABST
Patent Text Reader

Abstract

The invention discloses a container connection transportation path planning method based on deep reinforcement learning, and the method comprises the steps: building a mixed integer nonlinear programming model with the minimization of the total fuel consumption of all trucks as a target according to a container connection transportation system, performing linearization processing on the mixed integer nonlinear programming model to obtain a mixed integer linearization programming model; the method comprises the following steps: taking a container connection transportation path planning problem as a Markov decision process, and carrying out reinforcement learning training on a deep reinforcement learning network based on Transform in a container connection transportation simulation environment based on a deep reinforcement learning algorithm of Transform to obtain an optimal deep reinforcement learning network; and realizing container connection transportation path planning through the optimal deep reinforcement learning network. The method solves the problem that the existing method cannot effectively capture the structural features of the transportation process, limits the decision-making quality and generalization ability in high-dimensional and other complex scenes, and further causes the low container connection transportation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of container intermodal transport technology, and in particular to a container intermodal transport route planning method based on deep reinforcement learning. Background Technology

[0002] In the process of optimizing container relay transportation, it is necessary to first plan to meet various constraints (such as customer demand, time window, container type, truck capacity, etc.), then build a mathematical model of relay transportation, design an efficient solution algorithm suitable for the characteristics of the problem, and finally obtain an optimized relay transportation scheme to achieve effective utilization of transportation resources and reduction of relay transportation costs.

[0003] One of the main approaches to optimizing feeder transport is to build a mixed-integer programming model and then use commercial optimization software or intelligent optimization algorithms (such as large neighborhood search algorithms and simulated annealing algorithms) to obtain optimal or satisfactory solutions for different feeder transport scenarios, ultimately leading to an optimized feeder transport scheme regarding route planning and resource allocation. However, the intelligent optimization algorithms designed based on the mathematical model often struggle to effectively describe the complex spatial topological relationships between nodes (such as yards and customer points) in the feeder transport network, as well as the inherent symmetry and pairing relationships between pickup and delivery points. They also fail to effectively capture the structural characteristics of the transport process, limiting the decision-making quality and generalization ability of intelligent optimization algorithms in complex scenarios such as high dimensions, and ultimately resulting in low efficiency in container feeder transport. Summary of the Invention

[0004] This invention provides a container relay transportation route planning method based on deep reinforcement learning to overcome the above-mentioned technical problems.

[0005] To achieve the above objectives, the technical solution of the present invention is as follows: A method for planning container relay transportation routes based on deep reinforcement learning, specifically including the following steps: S1: Construct a container relay transportation system, which includes a set of request nodes for abstracting and describing the transportation scenario using graph theory techniques; the set of request nodes includes a subset of yard nodes for departure / return of container relay transportation, a subset of pickup nodes, and a subset of delivery nodes; S2: Based on the container relay transportation system, construct a mixed-integer nonlinear programming model with the objective of minimizing the total fuel consumption of all trucks; linearize the mixed-integer nonlinear programming model to obtain a mixed-integer linearized programming model; S3: Treat the container relay transportation route planning problem as a Markov decision process (MDP) and construct a Transformer-based deep reinforcement learning algorithm that includes an environment state space, action space, reward function, and a Transformer-based deep reinforcement learning network. The environmental state space includes the truck state and request node state used for container relay transportation; the request node state includes the request node coordinates, the container type of the request node task, the request node task type, and the access flag of the request node; the action space is the set space of the next request node that all trucks can choose; the reward function is used to define the feedback reward value, i.e., negative fuel consumption, based on the mixed integer linearized programming model to minimize the fuel consumption of the total path of all trucks. S4: A Transformer-based deep reinforcement learning algorithm is used to train a Transformer-based deep reinforcement learning network in a container relay transportation simulation environment based on the environment state space, action space, and reward function to obtain the optimal deep reinforcement learning network; the optimal deep reinforcement learning network is used to realize container relay transportation path planning.

[0006] Furthermore, the mixed-integer nonlinear programming model constructed in S2 includes mixed-integer nonlinear functions and mixed-integer nonlinear constraints: The expression for the mixed integer nonlinear function is:

[0007] In the formula: Indicates a collection of trucks. This represents the set of requesting nodes, including a subset of yard nodes and a subset of pickup nodes. and delivery node subset ,and , , ;0 indicates the starting and returning storage yard nodes; Indicates the number of requested tasks; Indicates the node that requests access to the truck. i The set of subsequent states; This indicates the fuel cost per kilometer for a truck. Indicates truck Access request node Decision variables for the state of the truck after it has been loaded; Indicates the request node With request node The driving distance between them; Indicates truck Whether to request node With request node Decision variables formed by the directed arc between them; The mixed-integer nonlinear constraint condition is: The constraint that all pickup nodes are visited only once:

[0008] In the formula: Represents a subset of pickup nodes; Constraint that any pickup node and its corresponding delivery node are served by the same truck:

[0009] In the formula: Indicates the number of requested tasks; Indicates truck Whether to request node With request node Decision variables formed by the directed arc between them; The constraint that each truck departs from and returns to the yard only once:

[0010] In the formula: Indicates truck Whether through the yard node With request node The decision variables formed by the directed arc between them. Indicates truck Whether to request node With the yard node Decision variables formed by the directed arc between them; The constraint that a truck arriving at a requesting node must leave that requesting node after completing its task:

[0011] If there is no state transition between two requesting nodes, the truck will not pass through the constraint of the directed arc between the two requesting nodes.

[0012] In the formula: This represents the set of directed arcs formed between all requesting nodes; This indicates that a truck passes through an arc in the set of directed arcs. The subsequent state transition pair; The constraint that trucks must access requesting nodes along the route in a predetermined order, i.e., the sequence of requesting nodes accessed by trucks must increase along the route:

[0013] In the formula: Indicates truck m Access pickup nodes along its travel route The order; Indicates truck m Access pickup nodes along its travel route The order; Represents positive integers; The constraint that trucks must access the pickup node before accessing the corresponding delivery node:

[0014] In the formula: Indicates truck m Access delivery nodes along its travel path +i The order; The constraint is that the sum of the order in which all trucks finally return to the yard node equals the number of nodes requested by all trucks.

[0015] In the formula: Indicates truck m The order in which they return to the yard node after completing their travel path; Indicates the number of trucks; If a truck does not visit a requested node, then that requested node cannot appear on the truck's route.

[0016] Constraints on the initial departure order of all trucks from the yard:

[0017] In the formula: Indicates truck m The initial order of departure from the storage yard; After a truck passes through a request node, only one truck state constraint is generated at that request node:

[0018] In the formula: Indicates truck Access Node The truck status after that is Decision variables; Truck status constraints in the yard:

[0019] In the formula: Indicates truck Status at the storage yard node; This represents the set of states of trucks at the yard nodes; Constraints on the change in the number of containers loaded after a truck access node executes its task:

[0020] In the formula: Indicates truck m Preparing to access the node i The number of containers previously loaded; Indicates truck m Preparing to access the node The number of containers previously loaded; Indicate design parameters and And if but ,otherwise ; Limits on the number of containers a truck can carry while in the yard:

[0021] In the formula: Indicates truck m The number of containers loaded at the yard; Truck Access Curve This will then result in a state transition, and the transitioned state will correspond to the truck's position on the arc. Constraints corresponding to the states of the two endpoints:

[0022]

[0023]

[0024] In the formula: Indicates if the truck Driving through the arc The subsequent state transition pair is The value is 1 if it is 1, otherwise it is 0. Indicates that the truck traveled across the arc The set of possible state transition pairs.

[0025] Furthermore, in S2, the mixed-integer nonlinear programming model is linearized to obtain the mixed-integer linearized programming model:

[0026]

[0027] In the formula: Indicates truck From the request node Fuel consumption at departure.

[0028] Furthermore, the reward function described in S3 The expression is:

[0029] In the formula: express The environment and state space at any given moment; express The action performed in the action space at a given moment; This represents the truck fuel consumption obtained based on a mixed-integer linearized programming model; express The number of containers loaded on the truck at any given time; Indicates the type of container.

[0030] Furthermore, the Transformer-based deep reinforcement learning network in S3 includes an input layer, an embedding layer, an encoder, and a decoder; The input layer is used to input preset transportation calculation examples into the embedding layer; Furthermore, the characteristic parameters of the transportation example include at least the truck state and request node state in the environmental state space, container type, number of trucks, and unit fuel consumption of the trucks; the label parameter of the transportation example is the selected request node. The embedding layer is used to perform node embedding processing on the request nodes in the transportation example parameters to obtain node embedding vectors, which include: Stockyard Node Embedding : Includes only stockpile node characteristics; Pickup node embedding : Includes only pickup node characteristics; Delivery node embedding : Includes only delivery node characteristics; Overall node embedding It is obtained by combining the characteristics of yard nodes, pickup nodes, and delivery nodes, and its expression is:

[0031]

[0032] In the formula: Represents the original feature matrix of the node; This represents a trainable weight tensor; Indicates a splicing operation; The encoder is used to extract the dependencies between yard nodes, pickup nodes, and delivery nodes based on a multi-head attention mechanism and node embedding vectors, obtaining the node dependency embedding vector, the expression of which is:

[0033]

[0034]

[0035]

[0036]

[0037]

[0038]

[0039]

[0040]

[0041]

[0042]

[0043]

[0044]

[0045] In the formula: res represents the reshaping operation; , , Represents a trainable weight tensor; Indicates the bias term; The number of heads indicating multi-head attention; This represents a query about a node corresponding to a certain attention head; Represents the key of a node corresponding to a certain attention head; This represents the value of a node corresponding to a certain attention head; Indicates the scaling factor. ; This represents dot product attention; Indicates attention; This represents the dimension of each attention head. Indicates the embedding dimension of the model; This represents the distribution of attention of the pickup node to all delivery nodes; This represents the distribution of attention of the delivery node to all pickup nodes; This represents the delivery node key corresponding to a certain attention head; This indicates a query for the pickup node corresponding to a specific attention point; This indicates the pickup node key corresponding to a specific attention point; This indicates a query for the delivery node corresponding to a specific attention head. This indicates the weight assigned to the attention level of each pickup node towards its corresponding delivery node; This indicates the weight assigned to the attention level of each delivery node towards its corresponding pickup node; This represents the dimension of the query, key, or value for each attention head; express The i dimension; express The i dimension; express The i dimension; express The i dimension; Represents the output tensor of multi-head self-attention; This represents the probability distribution that determines the attention allocation of the model across different inputs; , , , Represent the corresponding tensors respectively The original dot product attention that existed before the disassembled and restored connection operation; Represents the value vector of the delivery node; Represents the value vector of the pickup node; This represents the delivery node value vector after filling the tensor with zeros, and the padded dimension is the same as the tensor dimension. The dimensions are the same; This represents the vector of fetch node values ​​after filling the tensor with zeros. The dimensions after filling are the same as those of the tensor. The dimensions are the same; Indicates the bias term; , , Represents a trainable weight tensor; Indicates intermediate parameters; Represents the node dependency embedding vector and ; The decoder is used to decompose the node dependency embedding vector into dependency embeddings of the picking node. Dependency embedding with delivery nodes It also outputs the probability of the next request node being selected, based on the current environment state. The expression for this probability is:

[0046] In the formula: , Indicates intermediate parameters; , This represents the corresponding trainable weight tensor; Indicates a time index; This represents the corresponding trainable weight tensor; express The environment and state space at any given moment; , , Indicates the node dependency score; The mask vector representing the request node; , Indicates intermediate parameters; Indicates the probability that the next request node will be selected; This represents the attention score calculation result for each attention head in a multi-head attention mechanism. The result after combination; Indicates a constant parameter; This represents the activation function.

[0047] Furthermore, the Transformer-based deep reinforcement learning algorithm in S4, based on the environment state space, action space, and reward function, trains the Transformer-based deep reinforcement learning network in a container handling transportation simulation environment. The specific steps include: S41: Initialize the model policy parameters of the Transformer-based deep reinforcement learning network Set the maximum training period Batch size of the transportation example b Initialize the environment state space This includes the status of trucks used for container relay transportation. With the request node status Truck status ,in Indicates truck Drive from the yard to The length of the requested node at any given time; Indicated to The path that the truck travels at any given time is determined by... Time and corresponding selected request node Composition and , express The truck drove to the first One request node; Indicated to The number of containers loaded on the truck at any given time; , Indicates the maximum number of containers a truck can carry; the status of the requesting node. Including the requested node coordinates Container type of the request node task Request node task type and the access token of the request node ; S42: For each training cycle, input each transportation example into the Transformer-based deep reinforcement learning network, based on the current environment state. The probability of the next request node being selected is obtained by training a network model based on the parameters of the transportation case study. And based on the probability of being selected Confirm the next request node, i.e., the selected action in the action space. ; in performing the selected action Update environment status later That is, mark the visited request nodes and update the truck load; S43: Construct the advantage function for each transportation instance based on the reward function, and determine the probability of the next request node being selected for each transportation instance. The loss function for training the network model is:

[0048] =

[0049]

[0050]

[0051] In the formula: Indicates intermediate parameters; Represents the loss function; Representing a transportation example i Given network policy parameters At that time, the policy network is in Time and environmental state Selected action The probability of; Represents the dominance function; This represents the feedback reward for the result of each transportation example in a batch; Indicates the number of transport instances in the batch; Based on the probability of the next request node being selected Obtain the next selected actual prediction request node, and update the model policy parameters based on the loss function and the actual prediction request node using gradient descent. Repeat step S42 until the maximum training cycle is reached. And save the trained model policy parameters at this time. Obtain the optimal deep reinforcement learning network.

[0052] Beneficial Effects: This invention provides a container shuttle transportation route planning method based on deep reinforcement learning. It constructs a mixed-integer nonlinear programming model with the objective of minimizing the total fuel consumption of all trucks, and then linearizes it to obtain a mixed-integer linearized programming model. Compared to traditional methods that focus on travel distance or travel time, this invention effectively reduces transportation carbon emissions by minimizing fuel consumption, meeting the requirements of low-carbon transportation. Furthermore, it addresses multiple user access requests by splitting requests into virtual request nodes, and defines truck load states based on container type combinations to accurately adapt to mixed-size container scenarios. By treating the container shuttle transportation route planning problem as a Markov Decision Process (MDP) and constructing a Transformer-based deep reinforcement learning algorithm, it captures the network dependencies of the container shuttle transportation network through Transformer multi-head attention, overcoming the solution bottleneck of traditional intelligent optimization algorithms in large-scale scenarios. The Transformer-based deep reinforcement learning algorithm can also find the optimal solution in a short time and stably output effective and feasible solutions in large-scale scenarios, improving decision quality and generalization ability in complex scenarios such as high dimensions, thereby improving the efficiency of container shuttle transportation. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 This is a flowchart of the container intermodal transport route planning method based on deep reinforcement learning according to the present invention. Figure 2 This is the core block diagram of the deep reinforcement learning algorithm based on Transformer in this embodiment; Figure 3 This is a schematic diagram of the encoder structure in this embodiment; Figure 4 This is a schematic diagram of the decoder structure in this embodiment. Detailed Implementation

[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] This embodiment provides a container shuttle transportation route planning method based on deep reinforcement learning. Its main purpose is to overcome the shortcomings of existing vehicle route planning methods based on intelligent optimization algorithms and reinforcement learning, such as insufficient model expressive power, susceptibility to local optima, disconnect between optimization objectives and actual costs, and indirect decision-making patterns. Specifically, by integrating refined fuel cost modeling, an improved Transformer model architecture, and an end-to-end generative decision-making model, it aims to provide a new method capable of solving high-quality scheduling schemes that are both economical and low-carbon in large-scale, highly complex container shuttle transportation scenarios, thereby effectively addressing the deficiencies of existing technologies. Figure 1 As shown, the specific steps include: S1: Construct a container relay transportation system, which includes a set of request nodes for abstracting and describing the transportation scenario using graph theory techniques; the set of request nodes includes a subset of yard nodes for departure / return of container relay transportation, a subset of pickup nodes, and a subset of delivery nodes; Specifically, the actual connecting transportation scenario is abstractly described based on a "node-arc-state" graph, that is, graph theory techniques are used to abstractly describe the transportation scenario. It is assumed that the set of all request nodes to be processed is... Where 0 represents the departure and return yard nodes, and the remaining nodes correspond to customer nodes. This indicates the number of requested tasks, and the subset of pickup nodes is... The delivery node subset is For those belonging to a subset Each pickup node subset The corresponding transmission node is Assumption It is a directed graph. The set of directed arcs formed between nodes, and the set of task types. ,in This indicates a container pickup task at a specific customer location. This indicates a container delivery task at a specific customer location, with the container type set to... , These represent an empty 20-foot container and a full 20-foot container, respectively. Each customer node corresponds to a specific task type, container type, and container quantity. The customer node task executed by the truck is represented as follows: .in Represents a node The number of containers corresponding to the task.

[0057] S2: Based on the container relay transportation system, with the goal of minimizing the total fuel consumption of all trucks, a mixed-integer nonlinear programming model is constructed. The mixed-integer nonlinear programming model is then linearized to obtain a mixed-integer linearized programming model. Wherein, the total fuel consumption of all trucks = fuel consumption per unit distance per truck × travel distance. Specifically, the mixed-integer nonlinear programming model includes mixed-integer nonlinear functions and mixed-integer nonlinear constraints: The expression for the mixed integer nonlinear function is:

[0058] In the formula: Indicates a collection of trucks. This represents the set of requesting nodes, including a subset of yard nodes and a subset of pickup nodes. and delivery node subset ,and , , ;0 indicates the starting and returning storage yard nodes; Indicates the number of requested tasks; Indicates the node that requests access to the truck. i The set of subsequent states; This indicates the fuel cost per kilometer for a truck. Indicates truck Access request node The decision variables for the truck's state after it is in the following condition: The value is 1 if it is 1, otherwise it is 0. Indicates the request node With request node The driving distance between them; Indicates truck Whether to request node With request node If the decision variables are the directed arc formed between the trucks, then... via directed arc The value is 1 if it is 1, otherwise it is 0. The mixed-integer nonlinear constraint condition is: The constraint that all pickup nodes are visited only once:

[0059] In the formula: Represents a subset of pickup nodes; Constraint that any pickup node and its corresponding delivery node are served by the same truck:

[0060] In the formula: Indicates the number of requested tasks; Indicates truck Whether to request node With request node Decision variables formed by the directed arc between them; The constraint that each truck departs from and returns to the yard only once:

[0061] In the formula: Indicates truck Whether through the yard node With request node The decision variables formed by the directed arc between them. Indicates truck Whether to request node With the yard node Decision variables formed by the directed arc between them; The constraint that a truck arriving at a requesting node must leave that requesting node after completing its task:

[0062] If there is no state transition between two requesting nodes, the truck will not pass through the constraint of the directed arc between the two requesting nodes.

[0063] In the formula: This represents the set of directed arcs formed between all requesting nodes; This indicates that a truck passes through an arc in the set of directed arcs. The subsequent state transition pair; The constraint that trucks must access requesting nodes along the route in a predetermined order, i.e., the sequence of requesting nodes accessed by trucks must increase along the route:

[0064] In the formula: Indicates truck m Access pickup nodes along its travel route The order; Indicates truck m Access pickup nodes along its travel route The order; Represents positive integers; The constraint that trucks must access the pickup node before accessing the corresponding delivery node:

[0065] In the formula: Indicates truck m Access delivery nodes along its travel path +i The order; The constraint is that the sum of the order in which all trucks finally return to the yard node equals the number of nodes requested by all trucks.

[0066] In the formula: Indicates truck m The order in which they return to the yard node after completing their travel path; Indicates the number of trucks; If a truck does not visit a requested node, then that requested node cannot appear on the truck's route.

[0067] Constraints on the initial departure order of all trucks from the yard:

[0068] In the formula: Indicates truck m The initial order of departure from the storage yard is 1; After a truck passes through a request node, only one truck state constraint is generated at that request node:

[0069] In the formula: Indicates truck Access Node The truck status after that is The decision variable, if the state is The value is 1 if it is 1, otherwise it is 0. Truck status constraints in the yard:

[0070] In the formula: Indicates truck Status at the storage yard node; This represents the set of states of trucks at the yard nodes; Constraints on the change in the number of containers loaded after a truck access node executes its task:

[0071] In the formula: Indicates truck m Preparing to access the node i The number of containers previously loaded; Indicates truck m Preparing to access the node The number of containers previously loaded; Indicate design parameters and And if but ,otherwise ; Limits on the number of containers a truck can carry while in the yard:

[0072] In the formula: Indicates truck m The number of containers loaded at the yard; Truck Access Curve This will then result in a state transition, and the transitioned state will correspond to the truck's position on the arc. Constraints corresponding to the states of the two endpoints:

[0073]

[0074]

[0075] In the formula: Indicates if the truck Driving through the arc The subsequent state transition pair is The value is 1 if it is 1, otherwise it is 0. Indicates that the truck traveled across the arc The set of possible state transition pairs that may occur later; This indicates the maximum number of containers a truck can carry at one time; In a specific embodiment, the mixed-integer nonlinear programming model is linearized to obtain the mixed-integer linearized programming model:

[0076] Add linear constraints:

[0077] In the formula: Indicates truck From the request node Fuel consumption at departure; when hour, By minimizing the objective, we obtain In this case, constraints It is relaxed; when At that time, constraints It is relaxed, and Because minimizing the objective yields After the above linearization process, the original model is transformed into a mixed-integer linear programming model, which can be solved by... The optimal solution is obtained using software.

[0078] S3: The container shuttle transportation route planning problem is treated as a Markov Decision Process (MDP), and a Transformer-based deep reinforcement learning algorithm is constructed, including an environment state space, an action space, a reward function, and a Transformer-based deep reinforcement learning network. The environment state space includes the truck states and request node states for container shuttle transportation. The request node state includes the request node coordinates, the container type of the request node's task, the request node's task type, and the request node's access flag. The action space is the set space of the next request node that all trucks can choose. The reward function is used to define a feedback reward value, i.e., negative fuel consumption, based on a mixed-integer linearized programming model to minimize the total fuel consumption of all truck paths. Specifically, the container relay transportation route planning problem can be viewed as a Markov decision process. For large-scale scenarios (number of nodes > 17), a deep reinforcement learning algorithm based on Transformer is designed. The core of this algorithm is to overcome the computational bottleneck by using "multi-head attention to capture node dependencies + reinforcement learning dynamic decision-making." The algorithm architecture is as follows: Figure 1 As shown, the environmental state space in this embodiment This includes the status of trucks used for container relay transportation. With the request node status Truck status ,in Indicates truck Drive from the yard to The length of the requested node at any given time; Indicated to The path that the truck travels at any given time is determined by... Time and corresponding selected request node Composition and , express The truck drove to the first One request node; Indicated to The number of containers loaded on the truck at any given time; , Indicates the maximum number of containers a truck can carry; the status of the requesting node. Including the requested node coordinates Container type of the request node task Request node task type and the access token of the request node ;in, Two-dimensional coordinates; If the requesting node is a pickup task Otherwise, it is -1; if the requested node is accessed. Otherwise, it is 0; Action space is all possible actions. The set describes the set space to which all trucks can choose the next request node, and the actions. Indicates the truck is in The next service node selected at any given time, and the expression for the action space is:

[0079] This embodiment also includes action update rules: ① Add the node to the corresponding truck ② Update ,in Determined by the type of container mission; Reward function: The agent's goal is to maximize the long-term cumulative reward, i.e., minimize the total fuel consumption of all truck routes. In this embodiment, it will be... t The reward for a given moment is defined as negative fuel consumption:

[0080] In the formula: express The environment and state space at any given moment; express The action performed in the action space at a given moment; This represents the truck fuel consumption obtained based on a mixed-integer linearized programming model; express The number of containers loaded on the truck at any given time; This indicates the container type, and the corresponding maximum long-term cumulative reward is: .

[0081] In a specific embodiment, such as Figure 2 As shown, a Transformer-based deep reinforcement learning network includes an input layer, an embedding layer, an encoder, and a decoder. The input layer is used to input the preset transportation examples (i.e., transportation examples pre-generated based on actual needs and constraints) into the embedding layer; and the feature parameters of the transportation examples include at least the truck state and request node state in the environmental state space, container type, number of trucks, and unit fuel consumption of the trucks; the label parameter of the transportation examples is the selected request node; The embedding layer is used to perform node embedding processing on the request nodes in the transportation example parameters to obtain node embedding vectors, that is, to transform the "coordinates + task type + state constraints" of the request node into a high-dimensional vector, which includes: Stockyard Node Embedding : Includes only stockpile node characteristics; Pickup node embedding : Includes only pickup node characteristics; Delivery node embedding : Includes only delivery node characteristics; Overall node embedding It is obtained by combining the characteristics of yard nodes, pickup nodes, and delivery nodes, and its expression is:

[0082]

[0083] In the formula: Represents the original feature matrix of the node; This represents a trainable weight tensor; Indicates a splicing operation; In this embodiment, the encoder comprises multiple sub-layers, each consisting of "multi-head attention + feedforward neural network + residual connection", as shown in the following structure. Figure 3 As shown, the core is to capture three types of dependencies through multi-head attention. The encoder is used to extract the dependencies between yard nodes, pickup nodes, and delivery nodes based on the multi-head attention mechanism and the node embedding vectors, and obtain the node dependency embedding vectors. The specific steps include: S100: Acquiring multi-head attention at the pickup-delivery node is:

[0084]

[0085] In the formula: res represents the reshaping operation; , , Represents a trainable weight tensor; Indicates the bias term; The number of heads indicating multi-head attention; This represents a query about a node corresponding to a certain attention head; Represents the key of a node corresponding to a certain attention head; This represents the value of a node corresponding to a certain attention head; This represents a scaling factor, used to prevent the inner product in the dot product attention from becoming excessively large due to the increase in each head dimension. ; This represents dot product attention; Indicates attention; This represents the dimension of each attention head. Indicates the embedding dimension of the model; S101: Obtain the attention distribution of the pickup node with respect to all delivery nodes. The distribution of attention of delivery nodes to all pickup nodes for:

[0086]

[0087] In the formula: This represents the distribution of attention of the pickup node to all delivery nodes; This represents the distribution of attention of the delivery node to all pickup nodes; This represents the delivery node key corresponding to a certain attention head; This indicates a query for the pickup node corresponding to a specific attention point; This indicates the pickup node key corresponding to a specific attention point; This indicates a query for the delivery node corresponding to a specific attention head. S102: Obtain the attentional weights of each pickup node to its corresponding delivery node. And the weighting of attentional attention of each delivery node to its corresponding pickup node. for:

[0088]

[0089] In the formula: This indicates the weight assigned to the attention level of each pickup node towards its corresponding delivery node; This indicates the weight assigned to the attention level of each delivery node towards its corresponding pickup node; This represents the dimension of the query, key, or value for each attention head; express The i dimension; express The i dimension; express The i dimension; express The i dimension; Represents the output tensor of multi-head self-attention; S103: Based on steps S100 to S102, all attention heads are concatenated, and after normalization, reshaping, and linear transformation, the node dependency embedding vector is obtained as follows:

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096] In the formula: This represents the probability distribution that determines the attention allocation of the model across different inputs; , , , Represent the corresponding tensors respectively The original dot product attention that existed before the disassembled and restored connection operation; Represents the value vector of the delivery node; Represents the value vector of the pickup node; This represents the delivery node value vector after filling the tensor with zeros, and the padded dimension is the same as the tensor dimension. The dimensions are the same; This represents the vector of fetch node values ​​after filling the tensor with zeros. The dimensions after filling are the same as those of the tensor. The dimensions are the same; Indicates the bias term; , , Represents a trainable weight tensor; Indicates intermediate parameters; Represents the node dependency embedding vector and ; Represents the normalization function; This represents the activation function of the feedforward layer. In this embodiment, the final output shape and dimensions of the encoder sublayer are similar to... To maintain consistency, the output generated by the encoder sublayer will be used as the input data for the next sublayer in a cyclical manner, following the established calculation process described above, and the result output by the last sublayer of the entire encoder will be regarded as the overall output of the encoder and used in the subsequent decoder module.

[0097] The decoder is used to decompose the node dependency embedding vector into dependency embeddings of the picking node. Dependency embedding with delivery nodes The decoder, in this embodiment, outputs the probability of the next requested node being selected based on the encoder's output and the current environment state. The decoder structure is as follows: Figure 4 As shown. The number of heads, embeddings, queries, and key-value dimensions in the decoder are kept the same as those in the encoder, specifically including the following steps: S200: Output of encoder The current environment's state information is used as input to the decoder, where, The subscript 1 in the code is used to indicate the embedding information of the decoder, thus distinguishing it from the embedding in the encoder. Based on empirical values, the data is segmented to represent the embedding of the picking node. and delivery node embedded Then calculate the corresponding keys and values:

[0098]

[0099] In the formula: , This represents intermediate parameters, namely keys and values; , This represents the corresponding trainable weight tensor; S201: The decoder will embed the last selected node from the encoder's overall encoding based on the environmental state information recorded. Extract the embedding of the corresponding selected node and segment it into the embedding information of each head. Because multiple trucks will be used to complete the customer's task, the node embedding is calculated for each truck. That is, in this embodiment, the node embedding of the last selected node for each truck is calculated, rather than for a single request node. The expression is as follows:

[0100] In the formula: Indicates a time index; This represents the corresponding trainable weight tensor; S202: In the decoder, the symmetry between nodes is also considered, and the scores between the whole node and itself, the whole node to the pickup node, and the whole node to the delivery node are calculated respectively. for:

[0101] In the formula: express The environment and state space at any given moment; , , Indicates the node dependency score; The mask vector representing the request node; In this embodiment, after obtaining the information of the last node visited by the truck, the next node to be visited needs to exclude nodes that do not meet the constraints or cannot be visited according to the rules. The mask will be divided into a whole node mask. Pickup node mask and delivery node mask And it covers the scores from the overall node to the corresponding part of the request node, and its expression is:

[0102] S203: After concatenating and scaling the scores, normalization is performed. To avoid redundant calculations, only the scores of the entire node are calculated in slices, and then matrix multiplication is performed with the embedded values ​​of the entire node.

[0103] In the formula: Indicates to Perform a slicing operation to obtain values ​​from 0 to... Part of it; S204: Combine the results of the multi-head attention mechanism to obtain the overall computational score. ; S205: Calculate the overall score Embedded with the overall node After multiplication, an activation function is used to limit the range of the calculation result. Then, a mask is applied and normalization is performed to obtain the result. The probability of selecting the output node at that time is:

[0104] In the formula: Indicates the probability that the next request node will be selected; This represents the attention score calculation result for each attention head in a multi-head attention mechanism. The result after combination; Indicates a constant parameter; This represents the activation function.

[0105] S4: A Transformer-based deep reinforcement learning algorithm is used to train a Transformer-based deep reinforcement learning network in a container relay transportation simulation environment based on the environment state space, action space, and reward function to obtain the optimal deep reinforcement learning network; the optimal deep reinforcement learning network is used to realize container relay transportation path planning.

[0106] The method for obtaining the optimal deep reinforcement learning network in this embodiment is as follows: S41: Initialize the model policy parameters of the Transformer-based deep reinforcement learning network Set the maximum training period Batch size of the transportation example b and initialize the environment state; S42: For each training cycle, input each transportation example into the Transformer-based deep reinforcement learning network, based on the current environment state. The probability of the next request node being selected is obtained by training a network model based on the parameters of the transportation case study. And based on the probability of being selected Confirm the next request node, i.e., the selected action in the action space. ; in performing the selected action Update environment status later That is, mark the visited request nodes and update the truck load; S43: Construct the advantage function for each transportation instance based on the reward function, and determine the probability of the next request node being selected for each transportation instance. The loss function for training the network model is:

[0107] =

[0108]

[0109]

[0110] In the formula: Indicates intermediate parameters; Represents the loss function; Representing a transportation example i Given network policy parameters At that time, the policy network is in Time and environmental state Selected action The probability of; Represents the dominance function; This represents the feedback reward for the result of each transportation example in a batch; Indicates the number of transport instances in the batch; Based on the probability of the next request node being selected Obtain the next selected actual prediction request node, and update the model policy parameters based on the loss function and the actual prediction request node using gradient descent. Repeat step S42 until the maximum training cycle is reached. And save the trained model policy parameters at this time. Obtain the optimal deep reinforcement learning network.

[0111] Specifically, in this embodiment, the Deep Reinforcement Learning (DRL) algorithm of Transformer is divided into a training phase and a prediction phase: During the training phase, the policy gradient loss function is used to optimize the Transformer deep reinforcement learning network model, and its reward is defined as the negative of the total fuel consumption of the trucks used in a transportation example. The training steps are as follows: Step S001: Initialize the training model policy parameters The maximum training period is And the batch size of the transportation example is b ; Step S002: For each cycle: First generate b Each transportation case contains node coordinates, container type, and number of trucks; then, the environment state is initialized for each transportation case. s 0 (e.g., trucks are parked in the yard, and all nodes are not visited); Step S003: For each transportation example, generate a node dependency embedding vector through the encoder and Decoder output probability Select the node to be visited, and finally update the environment status. (For example: marking visited nodes, updating truck load); Step S004: After all transportation cases have been calculated, for each transportation case, the logarithmic probability of selecting each action (i.e., the next strong request node) is accumulated and summed to obtain... ;in, This indicates the transportation example. i Given model policy parameters At that time, the policy network is in Time and environmental state Select Action The probability of; Simultaneously calculate the dominance function for the results of each transportation example. in, This represents the feedback reward for the result of each transportation instance in a batch. Indicates the number of transport instances in the batch; Step S005: The loss function can be expressed as And update the model policy parameters through gradient descent. ; Step S006: Determine if the maximum training period has been reached. If not, return to step S002; otherwise, end training, update and save the model policy parameters. ; The prediction phase then begins, first loading the trained model policy parameters. Then, a new transportation example is input, and the node selection and environment state update process in the training process is repeated to obtain the optimized output results corresponding to the truck path and resource configuration.

[0112] This embodiment also includes system deployment on a workstation, with the following specific hardware configuration: the workstation is equipped with Intel(R) processors. (R) Gold 5218R CPU @ 2.10GHz, 128GB RAM, NVIDIA RTX A6000 graphics card, and a 64-bit operating system based on the x64 processor. For small-scale scenarios (no more than 17 request nodes), a solver is used. Version 10.0 solves the mathematical model in an environment built by Anaconda Virtual Environment Manager. The termination condition for the solution time of each example is set to 3600 seconds, or to terminate the run when the optimal solution is obtained. For large-scale scenarios (more than 17 request nodes), the DRL algorithm is used for solving.

[0113] The method described in this embodiment solves the core problem of optimizing urban container shuttle transportation through three core technological innovations: First, it handles multiple customer access requests by splitting the data into virtual nodes and defines truck load status by container type combination, which can accurately adapt to mixed-size container scenarios; second, it captures the network dependencies of container shuttle transportation based on Transformer multi-head attention, breaking through the solution bottleneck of traditional intelligent optimization algorithms in large-scale scenarios; third, it transforms the optimization objective from "minimizing driving distance" to "minimizing total fuel consumption," reconstructing the objective function and algorithm reward mechanism to achieve synergy between transportation efficiency and low-carbon requirements. Compared with existing technologies, the method described in this embodiment can achieve better performance in small-scale container shuttle scenarios by using a linearized model. The solver quickly finds the optimal solution, and the DRL algorithm can also find the optimal solution in a short time. In large-scale scenarios, the DRL algorithm can stably output effective and feasible solutions, demonstrating significant optimization effects and avoiding the undesirable results of traditional intelligent optimization algorithms, such as getting trapped in local optima. Furthermore, the method described in this embodiment aims to minimize fuel consumption, which, compared to traditional schemes that focus on travel distance or time, effectively reduces transportation carbon emissions and meets the requirements of low-carbon transportation. The method described in this embodiment can also effectively describe the complex spatial topology relationships between nodes (such as yards and customer points) in the connecting transportation network, as well as the inherent symmetry and pairing relationships between pickup and delivery nodes. It can effectively capture the structural characteristics of the transportation process, improving decision-making quality and generalization ability in complex scenarios such as high dimensions, thereby greatly improving the efficiency of container connecting transportation.

[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A container relay transportation route planning method based on deep reinforcement learning, characterized in that, The specific steps include: S1: Construct a container relay transportation system, which includes a set of request nodes for abstracting and describing the transportation scenario using graph theory techniques; the set of request nodes includes a subset of yard nodes for departure / return of container relay transportation, a subset of pickup nodes, and a subset of delivery nodes; S2: Based on the container relay transportation system, construct a mixed-integer nonlinear programming model with the objective of minimizing the total fuel consumption of all trucks; linearize the mixed-integer nonlinear programming model to obtain a mixed-integer linearized programming model; S3: Treat the container relay transportation route planning problem as a Markov decision process (MDP) and construct a Transformer-based deep reinforcement learning algorithm that includes an environment state space, action space, reward function, and a Transformer-based deep reinforcement learning network. The environmental state space includes the truck state and request node state used for container relay transportation; the request node state includes the request node coordinates, the container type of the request node task, the request node task type, and the access flag of the request node; the action space is the set space of the next request node that all trucks can choose; the reward function is used to define the feedback reward value, i.e., negative fuel consumption, based on the mixed integer linearized programming model to minimize the fuel consumption of the total path of all trucks. S4: A Transformer-based deep reinforcement learning algorithm is used to train a Transformer-based deep reinforcement learning network in a container relay transportation simulation environment based on the environment state space, action space, and reward function to obtain the optimal deep reinforcement learning network; the optimal deep reinforcement learning network is used to realize container relay transportation path planning.

2. The container relay transportation route planning method based on deep reinforcement learning according to claim 1, characterized in that, The mixed-integer nonlinear programming model constructed in S2 includes mixed-integer nonlinear functions and mixed-integer nonlinear constraints: The expression for the mixed integer nonlinear function is: In the formula: Indicates a collection of trucks. This represents the set of requesting nodes, including a subset of yard nodes and a subset of pickup nodes. and delivery node subset ,and , , ;0 indicates the starting and returning storage yard nodes; Indicates the number of requested tasks; Indicates the node that requests access to the truck. i The set of subsequent states; This indicates the fuel cost per kilometer for a truck. Indicates truck Access request node Decision variables for the state of the truck after it has been loaded; Indicates the request node With request node The driving distance between them; Indicates truck Whether to request node With request node Decision variables formed by the directed arc between them; The mixed-integer nonlinear constraint condition is: The constraint that all pickup nodes are visited only once: In the formula: Represents a subset of pickup nodes; Constraint that any pickup node and its corresponding delivery node are served by the same truck: In the formula: Indicates the number of requested tasks; Indicates truck Whether to request node With request node Decision variables formed by the directed arc between them; The constraint that each truck departs from and returns to the yard only once: In the formula: Indicates truck Whether through the yard node With request node The decision variables formed by the directed arc between them. Indicates truck Whether to request node With the yard node Decision variables formed by the directed arc between them; The constraint that a truck arriving at a requesting node must leave that requesting node after completing its task: If there is no state transition between two requesting nodes, the truck will not pass through the constraint of the directed arc between the two requesting nodes. In the formula: This represents the set of directed arcs formed between all requesting nodes; This indicates that a truck passes through an arc in the set of directed arcs. The subsequent state transition pair; The constraint that trucks must access requesting nodes along the route in a predetermined order, i.e., the sequence of requesting nodes accessed by trucks must increase along the route: In the formula: Indicates truck m Access pickup nodes along its travel route The order; Indicates truck m Access pickup nodes along its travel route The order; Represents positive integers; The constraint that trucks must access the pickup node before accessing the corresponding delivery node: In the formula: Indicates truck m Access delivery nodes along its travel path +i The order; The constraint is that the sum of the order in which all trucks finally return to the yard node equals the number of nodes requested by all trucks. In the formula: Indicates truck m The order in which they return to the yard node after completing their travel path; Indicates the number of trucks; If a truck does not visit a requesting node, then that requesting node cannot appear on the truck's route. Constraints on the initial departure order of all trucks from the yard: In the formula: Indicates truck m The initial order of departure from the storage yard; After a truck passes through a request node, only one truck state constraint is generated at that request node: In the formula: Indicates truck Access Node The truck status after that is Decision variables; Truck status constraints in the yard: In the formula: Indicates truck Status at the storage yard node; This represents the set of states of trucks at the yard nodes; Constraints on the change in the number of containers loaded after a truck access node executes its task: In the formula: Indicates truck m Preparing to access the node i The number of containers previously loaded; Indicates truck m Preparing to access the node The number of containers previously loaded; Indicate design parameters and And if but ,otherwise ; Limits on the number of containers a truck can carry while in the yard: In the formula: Indicates truck m The number of containers loaded at the yard; Truck Access Curve This will then result in a state transition, and the transitioned state will correspond to the truck's position on the arc. Constraints corresponding to the states of the two endpoints: In the formula: Indicates if the truck Driving through the arc The subsequent state transition pair is The value is 1 if it is 1, otherwise it is 0. Indicates that the truck traveled across the arc The set of possible state transition pairs that may occur later.

3. The container relay transportation route planning method based on deep reinforcement learning according to claim 2, characterized in that, In S2, the mixed-integer nonlinear programming model is linearized to obtain the mixed-integer linearized programming model as follows: In the formula: Indicates truck From the request node Fuel consumption at departure.

4. The container relay transportation route planning method based on deep reinforcement learning according to claim 3, characterized in that, The reward function described in S3 The expression is: In the formula: express The environment and state space at any given moment; express The action performed in the action space at a given moment; This represents the truck fuel consumption obtained based on a mixed-integer linearized programming model. express The number of containers loaded on the truck at any given time; Indicates the container type.

5. The container relay transportation route planning method based on deep reinforcement learning according to claim 4, characterized in that, The Transformer-based deep reinforcement learning network in S3 includes an input layer, an embedding layer, an encoder, and a decoder. The input layer is used to input preset transportation calculation examples into the embedding layer; Furthermore, the characteristic parameters of the transportation example include at least the truck state and request node state in the environmental state space, container type, number of trucks, and unit fuel consumption of the trucks; The label parameter of the transportation example is the selected request node; The embedding layer is used to perform node embedding processing on the request nodes in the transportation example parameters to obtain node embedding vectors, which include: Stockyard Node Embedding : Includes only yard node characteristics; Pickup node embedding : Includes only pickup node characteristics; Delivery node embedding : Includes only delivery node characteristics; Overall node embedding It is obtained by combining the characteristics of yard nodes, pickup nodes, and delivery nodes, and its expression is: In the formula: Represents the original feature matrix of the node; This represents a trainable weight tensor; Indicates a splicing operation; The encoder is used to extract the dependencies between yard nodes, pickup nodes, and delivery nodes based on a multi-head attention mechanism and node embedding vectors, obtaining the node dependency embedding vector, the expression of which is: In the formula: res represents the reshaping operation; , , Represents a trainable weight tensor; Indicates the bias term; The number of heads indicating multi-head attention; This represents a query about a node corresponding to a certain attention head; Represents the key of a node corresponding to a certain attention head; This represents the value of a node corresponding to a certain attention head; Indicates the scaling factor. ; This represents dot product attention; Indicates attention; This represents the dimension of each attention head. Indicates the embedding dimension of the model; This represents the distribution of attention of the pickup node to all delivery nodes; This represents the distribution of attention of the delivery node to all pickup nodes; This represents the delivery node key corresponding to a certain attention head; This indicates a query for the pickup node corresponding to a specific attention point; This indicates the pickup node key corresponding to a specific attention point; This indicates a query for the delivery node corresponding to a specific attention head. This indicates the weight assigned to the attention level of each pickup node towards its corresponding delivery node; This indicates the weight assigned to the attention level of each delivery node towards its corresponding pickup node; This represents the dimension of the query, key, or value for each attention head; express The i dimension; express The i dimension; express The i dimension; express The i dimension; Represents the output tensor of multi-head self-attention; This represents the probability distribution that determines the attention allocation of the model across different inputs; , , , Represent the corresponding tensors respectively The original dot product attention that existed before the disassembled and restored connection operation; Represents the value vector of the delivery node; Represents the value vector of the pickup node; This represents the delivery node value vector after filling the tensor with zeros. The dimensions after filling are equal to the tensor's dimensions. The dimensions are the same; This represents the vector of retrieved node values ​​after filling the tensor with zeros. The dimensions after filling are the same as those of the tensor. The dimensions are the same; Indicates the bias term; , , Represents a trainable weight tensor; Indicates intermediate parameters; Represents the node dependency embedding vector and ; The decoder is used to decompose the node dependency embedding vector into dependency embeddings of the picking node. Dependency embedding with delivery nodes It also outputs the probability of the next request node being selected, based on the current environment state. The expression for this probability is: In the formula: , Indicates intermediate parameters; , This represents the corresponding trainable weight tensor; Indicates a time index; This represents the corresponding trainable weight tensor; express The environment and state space at any given moment; , , Indicates the node dependency score; The mask vector representing the request node; , Indicates intermediate parameters; Indicates the probability that the next request node will be selected; This represents the attention score calculation result for each attention head in a multi-head attention mechanism. The result after combination; Indicates a constant parameter; This represents the activation function.

6. The container relay transportation route planning method based on deep reinforcement learning according to claim 5, characterized in that, The Transformer-based deep reinforcement learning algorithm in S4 describes a method for training a Transformer-based deep reinforcement learning network in a container handling transportation simulation environment, based on the environment's state space, action space, and reward function. The specific steps include: S41: Initialize the model policy parameters of the Transformer-based deep reinforcement learning network Set the maximum training period Batch size of transportation example b Initialize the environment state space This includes the status of trucks used for container relay transportation. With the request node status Truck status ,in Indicates truck Drive from the yard to The length of the requested node at any given time; Indicated to The path that the truck travels at any given time is determined by... Time and corresponding selected request node Composition and , express The truck drove to the first One request node; Indicated to The number of containers loaded on the truck at any given time; , Indicates the maximum number of containers a truck can carry; The request node status Including the requested node coordinates Container type of the request node task Request node task type and the access token of the request node ; S42: For each training cycle, input each transportation example into the Transformer-based deep reinforcement learning network, based on the current environment state. The probability of the next request node being selected is obtained by training a network model based on the parameters of the transportation case study. And based on the probability of being selected Confirm the next request node, i.e., the selected action in the action space. ; Execute the selected action Update environment status later That is, mark the visited request nodes and update the truck load; S43: Construct the advantage function for each transportation instance based on the reward function, and determine the probability of the next request node being selected for each transportation instance. The loss function for training the network model is: = In the formula: Indicates intermediate parameters; Represents the loss function; Representing a transportation example i Given network policy parameters At that time, the policy network is in Time and environmental state Selected action The probability of; Represents the dominance function; This represents the feedback reward for the result of each transportation example in a batch; Indicates the number of transport instances in the batch; Based on the probability of the next request node being selected Obtain the next selected actual prediction request node, and update the model policy parameters based on the loss function and the actual prediction request node using gradient descent. Repeat step S42 until the maximum training cycle is reached. And save the trained model policy parameters at this time. Obtain the optimal deep reinforcement learning network.