Logistics distribution vehicle intelligent scheduling method, system and device based on space-time attention mechanism, and medium

By using a deep reinforcement learning algorithm based on a spatiotemporal attention mechanism, the problem of complex optimization requirements in multi-vehicle scheduling is solved, and efficient vehicle route planning under practical constraints is achieved, improving transportation efficiency and accuracy.

CN120931004AActive Publication Date: 2025-11-11HARBIN INST OF TECH AT WEIHAI

Patent Information

Application Number
CN202511055128.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-11
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

Existing technologies are insufficient to meet the complex optimization needs of large-scale, high-dimensional scenarios in multi-vehicle scheduling. They neglect timeout costs and fail to fully consider the actual constraints in the cargo transportation process, leading to problems in the execution of the generated scheduling scheme.

Method used

A deep reinforcement learning algorithm based on spatiotemporal attention mechanism is adopted. The temporal and spatial features of orders are transformed into feature vectors through the embedding layer. Spatial and temporal attention mechanisms are used for processing. Combined with gating fusion mechanism, vehicle route planning is generated. The target parameter network is trained through reinforcement learning to ensure that the model optimizes the scheduling scheme under actual constraints.

Benefits of technology

This reduces the total route length of vehicles, decreases the occurrence of delays, increases vehicle load factor and speed, improves the transportation efficiency of express delivery vehicles, shortens delivery time, and ensures the rationality and feasibility of the generated scheduling plan in actual implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931004A_ABST
    Figure CN120931004A_ABST
Patent Text Reader

Abstract

The invention provides a logistics distribution vehicle intelligent scheduling method, system and device based on a space-time attention mechanism, and a medium, belongs to the technical field of artificial intelligence and logistics transportation, and specifically relates to collection of an express order set and construction of an express vehicle transportation network diagram; constructing an express vehicle scheduling model, and setting constraint conditions; designing a deep reinforcement learning algorithm based on a space-time attention mechanism, and carrying out order distribution and vehicle route planning; constructing a reinforcement learning algorithm based on actors and commentators, and training a target parameter network; and generating a vehicle route, calculating a reward value of the route and a state value of the valuation network, and completing training of a preset round based on updating parameters of the actor strategy network and the commentator network to obtain a target parameter network. According to the invention, the total path length of vehicle driving is reduced, the intersection and repetition of the paths are avoided, and the load factor and the driving speed of the vehicle are improved, so that the transportation efficiency of the express vehicle is improved, and the express delivery time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and logistics transportation technology, specifically relating to a method, system, equipment and medium for intelligent scheduling of logistics delivery vehicles based on a spatiotemporal attention mechanism. Background Technology

[0002] With the rapid development of the express delivery and urban logistics industries, logistics transportation systems face a series of complex challenges, including ever-increasing order demand, higher delivery time requirements, and limited transportation resources. In typical multi-vehicle dispatching scenarios, the system needs to rationally allocate multiple transportation tasks among limited fleet resources to minimize transportation costs, improve service efficiency, and meet customer needs.

[0003] The vehicle routing and scheduling methods in related technologies mainly rely on heuristic algorithms (such as nearest neighbor, ant colony algorithm, and genetic algorithm) and exact algorithms (such as integer linear programming and branch and bound). Although these methods perform reasonably well in static problems, their solution efficiency is not high, making it difficult to meet the complex optimization needs of large-scale, high-dimensional scenarios.

[0004] The related technologies aim to minimize the length of transportation routes, ignoring the impact of overtime costs on overall operations. Furthermore, the constraints are simple, potentially only considering vehicle weight limits and failing to adequately account for crucial operational constraints such as preventing loops in the same vehicle route during transport. This can lead to various problems in the actual execution of the generated scheduling schemes, making them unable to meet complex express delivery needs.

[0005] Order allocation and vehicle route planning methods often consider time and spatial characteristics separately, planning based on the geographical location of delivery and pickup nodes without analyzing the regional distribution patterns of nodes or the actual conditions of the regional traffic network. This leads to unreasonable order allocation and vehicle route planning that cannot effectively adapt to actual time and space constraints, easily resulting in vehicles being over-concentrated in certain areas or failing to complete tasks on time due to improper time scheduling. Summary of the Invention

[0006] This invention provides an intelligent scheduling method for logistics delivery vehicles based on a spatiotemporal attention mechanism. From data collection, constructing a reasonable model, and accurately planning routes to stable learning and optimization strategies, it reduces the total path length of vehicle travel, reduces the occurrence of timeouts, avoids route intersections and repetitions, and improves the vehicle's load factor and travel speed, thereby improving the transportation efficiency of express delivery vehicles and shortening the express delivery time.

[0007] The methods include: S101: Collect a set of express delivery orders, construct a transportation network diagram of express delivery vehicles, and obtain basic information about the set of transportation vehicles; S102: Construct a delivery vehicle scheduling model, setting the model objective as minimizing the sum of the total transportation route length and overtime cost of all vehicles, and setting constraints. S103: Design a deep reinforcement learning algorithm based on spatiotemporal attention mechanism for order allocation and vehicle route planning; wherein, the deep reinforcement learning algorithm transforms the temporal and spatial features of the order into feature vectors through an embedding layer, and inputs them into the spatial attention mechanism and temporal attention mechanism for processing respectively. After a preset number of iterations, a customer vector is generated through a gating fusion mechanism, and then the access order of each vehicle is determined through policy decoding to complete order allocation and vehicle route planning for each route from the warehouse to the warehouse after completing the task; S104: Construct a reinforcement learning algorithm based on actor and critic networks to train the target parameter network; wherein, the reinforcement learning algorithm initializes the actor policy network and the critic network, and through multiple rounds of training, calculates the comprehensive embedding features of customers in each round, generates vehicle routes, calculates the reward value of the routes and the state value of the valuation network, updates the parameters of the actor policy network and the critic network based on this, completes the training for the preset number of rounds, and obtains the target parameter network.

[0008] According to another embodiment of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the intelligent scheduling method for logistics delivery vehicles based on a spatiotemporal attention mechanism.

[0009] According to another embodiment of this application, a storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the intelligent scheduling method for logistics delivery vehicles based on a spatiotemporal attention mechanism.

[0010] As can be seen from the above technical solutions, the present invention has the following advantages: The intelligent scheduling method for logistics delivery vehicles based on a spatiotemporal attention mechanism provided by this invention collects order sets, covering detailed information such as order number, pickup / delivery type, customer location, cargo weight and volume, and time window. When constructing the transportation network map, it accurately marks the geographical information of warehouses, delivery and pickup nodes, and determines the optimal route distance between nodes considering actual road traffic conditions. At the same time, it acquires comprehensive vehicle information such as vehicle number, vehicle type, load capacity, driving speed, fuel consumption rate, and driving range, ensuring that all elements involved in scheduling are accurately quantified and represented. This lays a reliable data foundation for the entire scheduling process, enabling subsequent model construction and route planning to be based on real and complete data, avoiding unreasonable scheduling schemes due to missing or inaccurate information.

[0011] This invention sets the model objective as minimizing the sum of the total transportation path length and overtime cost for all vehicles. By incorporating these two key cost factors in the transportation process into the optimization scope, the scheduling scheme is designed from the outset to reduce overall costs, thereby improving the utilization efficiency of transportation resources. In addition to basic access and driving balance constraints, it also includes detailed constraints such as vehicle load not exceeding maximum load capacity, goods not being damaged during transportation, and the absence of loops in the routes of the same vehicle. These ensure the feasibility and rationality of the scheduling scheme and prevent situations where execution is impossible due to violations of actual constraints.

[0012] This invention strengthens the representation of spatiotemporal features through a preset number of iterations. The gating fusion mechanism dynamically allocates weights based on the impact of spatiotemporal features on the current scheduling task, generating a comprehensive customer vector. During strategy decoding, the access order is adjusted in real time based on vehicle remaining load, travel time, and other statuses, reducing route intersections and duplications, improving vehicle transportation efficiency, and adapting to the time and space constraints of different orders, thereby enhancing the flexibility and accuracy of scheduling.

[0013] During network initialization, network parameters are pre-adjusted based on historical scheduling data to give the initial network a certain scheduling capability. In multiple training rounds, reward value calculation comprehensively considers factors such as transportation distance, timeout cost, vehicle load factor, and route smoothness. Gradient pruning techniques are used during parameter updates to prevent gradient explosion, ensuring stable network updates. This allows the target parameter network to continuously learn and optimize scheduling strategies, adapting to different orders and transportation scenarios. The preset number of rounds is dynamically adjusted based on changes in network performance. When the routes generated by the network perform stably and excellently in multiple tests, training can be terminated early, improving training efficiency while ensuring a better generated scheduling scheme, thus enhancing the practicality and reliability of the scheduling model. Attached Figure Description

[0014] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 The flowchart shows a method for intelligent scheduling of logistics delivery vehicles based on a spatiotemporal attention mechanism. Figure 2 This is a schematic diagram of the express delivery vehicle transportation scheduling of the present invention; Figure 3 This is a diagram of the deep reinforcement learning model based on spatiotemporal attention mechanism of the present invention; Figure 4 This is a flowchart of the deep reinforcement learning process based on the spatiotemporal attention mechanism of the present invention. Figure 5 This is a flowchart illustrating the training process of the reinforcement learning algorithm based on actor critics according to the present invention. Figure 6 This is a schematic diagram of an electronic device. Detailed Implementation

[0016] The intelligent scheduling method for logistics delivery vehicles based on spatiotemporal attention mechanism provided by this invention includes: collecting order sets, constructing a transportation network graph for express delivery vehicles, and obtaining basic information about the transportation vehicle set; constructing an express delivery vehicle scheduling model; designing a deep reinforcement learning algorithm based on spatiotemporal attention to allocate orders and plan vehicle routes; and constructing a target parameter network based on an actor / critic reinforcement learning algorithm. This method can alleviate the delivery pressure on express delivery outlets or last-mile stations and improve delivery efficiency.

[0017] The following describes in detail the intelligent scheduling method for logistics delivery vehicles based on a spatiotemporal attention mechanism according to this application. Specific details, such as particular system structures and technologies, are presented for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application can also be implemented in other embodiments without these specific details.

[0018] The statements such as "one embodiment" or "some embodiments" described in this application mean that one or more embodiments of this application include the specific features, structures, or characteristics described in that embodiment. Therefore, the statements such as "in one embodiment," "in some embodiments," "in other embodiments," and "in still other embodiments" in this application do not necessarily refer to the same embodiment, but rather mean one or more, but not all, embodiments, unless otherwise specifically emphasized.

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please see Figure 1 The diagram shows a flowchart of a logistics delivery vehicle intelligent scheduling method based on a spatiotemporal attention mechanism in a specific embodiment. The method includes: Step S101: Collect a set of express delivery orders, construct an express delivery vehicle transportation network graph, and obtain basic information about the transportation vehicle set; wherein, the transportation network graph includes warehouse nodes, delivery nodes, pickup nodes, and edges between nodes, the edges represent the optimal routes and distances between nodes, the distances satisfy symmetry and the distance between the same node is 0.

[0021] In some embodiments, when collecting express delivery orders, it is necessary to extract key information for each order, including pickup / delivery time windows, such as the earliest pickup time and the latest delivery time; cargo weight / volume, used to match vehicle load capacity; geographic coordinates of pickup and delivery points; when constructing the transportation network map, the optimal route distance between each node is calculated using a Geographic Information System (GIS) or the Dijkstra algorithm to ensure that the distance of the edges satisfies symmetry, that is, the distance from i to j is equal to the distance from j to i, and connections are established between warehouse nodes and other nodes; when obtaining basic vehicle information, it is necessary to record the maximum load capacity, maximum daily mileage, and location of the express delivery station to which each vehicle belongs.

[0022] In some specific embodiments, step S101 specifically includes: Step 1.1: Construct a transportation network diagram for express delivery vehicles , where V contains Warehouse node, delivery node and pickup nodes , This represents the set of delivery nodes and pickup nodes. For the set of edges between nodes, each edge This represents the optimal route from node i to node j. Let represent the distance to edge (i,j), where the distance satisfies symmetry, i.e. The distance between the same nodes is 0, that is , .

[0023] Step 1.2: Indicates the set of available vehicles. x For the number of vehicles, Indicates vehicle k The maximum load capacity of the vehicle is determined by the courier station, which is the starting point for the vehicle and its final destination.

[0024] Appendix Figure 2This diagram illustrates the dispatching of delivery vehicles. In this scenario, all delivery tasks originate from and terminate at the central warehouse, involving multiple pickup and delivery task nodes. The system dispatches two delivery vehicles (Vehicle 1 and Vehicle 2), each undertaking different transportation tasks. Vehicle 1 (green) executes Route 1 (black path), sequentially proceeding to nodes in the upper area to perform tasks, including multiple pickup and delivery operations. Vehicle 2 (white) executes Route 2 (red path), primarily covering the lower area, completing the pickup and delivery tasks for another portion of orders. Each route returns to the central warehouse after completing all tasks. During the dispatching process, the system rationally plans the access order of each vehicle based on the geographical distribution of each node, task type (pickup or delivery), and timing requirements, minimizing route overlap and duplication while improving vehicle transportation efficiency and resource utilization.

[0025] This embodiment addresses the problem in existing technologies where nodes in express delivery networks often lack clear distinction between delivery and pickup nodes. Edge distances are frequently simplified to straight-line distances or symmetry is ignored, leading to chaotic spatial features that are difficult for attention mechanisms to capture effectively. Step 1.1 precisely defines node types, such as warehouses, delivery nodes (D), and pickup nodes (P), using the graph G=(V,E). It calculates the optimal route distance for edges and clarifies distance symmetry and the property that self-distance is zero, giving the spatial structure of the transportation network a rigorous mathematical logic. This structured modeling provides resolvable spatial features for the spatial attention mechanism in step S103. For example, spatial attention can calculate association weights based on the clustering of node types and simplify the computational complexity of cross-node spatial associations based on distance symmetry. This solves the problem in existing technologies where fuzzy spatial features make it difficult for attention mechanisms to accurately capture node associations, thus improving the ability of spatial attention to model the topology of transportation networks.

[0026] Step 1.2 defines the maximum load capacity of vehicle set K and the closed-loop constraint from departure to return from the station, providing key dynamic constraints for the time attention mechanism in step S103: on the one hand, the vehicle load capacity limits the total number of tasks in a single transport, and time attention needs to adjust the task sequence accordingly; on the other hand, the round-trip constraint requires time attention to ensure route closure when planning the timing, so as to avoid task scheduling exceeding the vehicle's range or time window.

[0027] The node set V and edge set E defined in step 1.1 provide a clear target for the node access constraints in step S102. For example, constraining each delivery node... Accessed only once can be directly mapped to the classification of V; the distance of the edges becomes the basis for calculating the total transportation path length in the objective function, transforming the goal of minimizing the total path from abstract to based on the actual optimal route distance. The vehicle load and round-trip constraints in step 1.2 are then transformed into specific parameters for the load constraints and closed-loop constraints in step S102, ensuring that the model constraints do not deviate from the actual vehicle capabilities. This fusion transforms the constraints of the scheduling model from general rules to personalized rules bound to specific network and vehicle attributes, breaking through the limitation of existing technologies where model constraints are disconnected from the actual scenario, and its non-obviousness is reflected in...

[0028] In step 1.1, the node classification and distance are directly used as the raw input for spatial feature extraction in step S103. When calculating node association weights, the spatial attention mechanism prioritizes strengthening the association between delivery nodes and pickup nodes within the same area, serving as the basic measure of spatial similarity. The symmetry of edges simplifies the bidirectional calculation of attention weights, avoiding redundancy. The vehicle load and round-trip constraints in step 1.2 are integrated into the temporal attention mechanism in step S103 through feature embedding: the difference between the vehicle's current load and its remaining load serves as a time urgency feature; for example, the smaller the remaining load, the stricter the time window constraints for subsequent tasks. The round-trip constraint serves as a route closure feature, ensuring that the time association weight between the last task node and the warehouse node is maximized in the iterative calculation of temporal attention. This fusion allows the spatiotemporal attention mechanism to no longer rely on general features, but rather dynamically adjust attention weights based on the specific attributes of the transportation network and vehicles, solving the problem of excessive generalization and insufficient scenario adaptation in existing attention mechanisms.

[0029] Step S102: Construct a delivery vehicle scheduling model, setting the model objective as minimizing the sum of the total transportation route length and overtime cost of all vehicles, and setting constraints.

[0030] In some embodiments, the model objective function needs to quantify the total transportation cost, specifically including two parts: the total transportation path length, which can be the sum of the actual distances of all vehicle routes; and timeout costs. Constraints cover business rules, such as: each order can only be accessed by one vehicle once; vehicles depart from and return from their respective warehouses; the total load capacity of a vehicle at any given time cannot exceed its maximum load capacity; and the arrival time of a vehicle at each node cannot be earlier than the start time of that node's time window and cannot be later than the end time. It can be seen that mathematical modeling transforms the scheduling problem into a constrained optimization problem. The objective function defines the optimal evaluation criteria, and the constraints limit the legal solution space; both together guide the algorithm to search for a reasonable scheduling scheme.

[0031] In some specific embodiments, step S102 defines Indicates vehicle From point Drive directly to the destination If so, then ,otherwise .

[0032] The optimal scheduling model for express delivery vehicles is established as follows:

[0033] In the formal model above, equation (1) is the objective function, representing minimizing the total transportation route and overtime cost of all vehicles; equation (2) indicates that each customer node is visited only once; equation (3) indicates that if the vehicle k When a vehicle visits a customer, it must also leave that customer; Equation (4) indicates that the vehicle will definitely depart from the warehouse and return to the warehouse; Equations (5) and (6) indicate that after satisfying the customer's needs, the vehicle... k The capacity is adjusted accordingly; Equation (7) indicates that vehicle k will carry the demand of all delivery nodes when it departs from the warehouse; Equation (8) indicates that the vehicle capacity cannot be exceeded by the total demand on a single path; Equation (9) indicates that if the vehicle k From node i To the node j His arrival time must be greater than the node. i Arrival time plus node i The service time is expressed by Equation (10), which indicates that the service time of a node is no earlier than the earliest service time window of that node.

[0034] In the formula, M is a positive number used to transform logical constraints (such as 0-1 variable constraints on whether a vehicle visits a node) into linear constraints. When vehicle k does not visit node i, M The condition takes effect to enforce the constraint; when vehicle k accesses node i, M The disappearance of an item does not affect the constraint.

[0035] β is the weighting coefficient for timeout costs. The objective function minimizes the sum of the total transportation route length and the timeout cost. β Used to balance the relative importance of the two, such as β The larger the value, the heavier the penalty for exceeding the time limit on the overall goal. is the time when vehicle k arrives at node j. represents the timestamp after vehicle k completes its service to node j. Let be the deadline for the time window of node j. This represents the latest time that node j requires vehicle k to arrive and complete the service.

[0036] j , i The nodes are indexes, all belonging to the node set. V Includes repository nodes V 0. Delivery Nodes D Pick-up nodes P . iand j Represents any two nodes, used to describe the relationship between the nodes.

[0037] is the time when vehicle k arrives at node i. represents the timestamp after vehicle k completes its service to node i. Let be the travel time of vehicle k from node i to node j. Let represent the time consumed by the vehicle on the journey from node i to node j. Let K be the remaining load capacity of vehicle K after the operation at node J. Let be the remaining load capacity of vehicle k after the operation at node i. Let K be the initial load capacity of vehicle k when it departs from the warehouse, or the remaining load capacity after it returns to the warehouse. Let be the demand for goods at node i, where picking is positive and delivery is negative. If node i is a picking node, di >0 requires loading goods; if it is a delivery node. di <0 requires unloading of goods. is the service time for node i. represents the time required for vehicle k to perform loading and unloading operations at node i, such as unloading time and loading time. Let be the service time for vehicle k from node i to node j. Let be the earliest service time window for node i. Let represent the earliest allowed time for vehicle k to arrive at node i.

[0038] In this embodiment, the objective function of Equation 1 in step S102 combines the total transportation path length with the timeout cost, injecting a clear cost orientation into the spatiotemporal attention mechanism: when calculating the node association weights, spatial attention prioritizes strengthening the association of short-path node pairs to reduce the total path length; when planning the time sequence, temporal attention increases the priority of nodes with ample time windows. This design, which binds abstract attention weights with concrete cost objectives, ensures that feature learning is no longer divorced from actual operating costs, resolving the contradiction in existing technologies that optimal features do not necessarily equal optimal costs.

[0039] The total path cost and timeout cost in equation (1) are used as penalty factors for spatial / temporal attention. When calculating the spatial attention weights of nodes i and j, if the path distance is too large, the weights are automatically reduced; when calculating the temporal attention weights, if the time window of j is tight, the weights are increased. Constraints (2)-(10) are transformed into a mask matrix. In the softmax calculation of spatiotemporal attention, node pairs that violate the constraints, such as visited nodes, overloaded paths, and nodes with reversed time order, are given extremely low weights, so that the model naturally does not pay attention to infeasible associations.

[0040] In step S104, the route reward value is directly calculated based on equation (1) (reward = -(total path cost + timeout cost)), allowing the actor network to naturally evolve towards the model optimization goal when generating routes. For routes that violate constraints (such as repeated node visits, overload), an additional penalty is added to the reward value, such as deducting 10% of the reward for violating constraint (2), so that the critic network prioritizes rejecting infeasible routes when evaluating state value. In the policy gradient calculation, higher gradient weights are assigned to paths that satisfy the constraints, and lower weights are assigned to paths that violate the constraints. This fusion makes the reinforcement learning training process and the optimization goal and constraints of the scheduling model form a closed loop, ensuring that the routes generated by the trained target parameter network not only meet the model requirements but also have practical feasibility.

[0041] Step S103: Design a deep reinforcement learning algorithm based on spatiotemporal attention mechanism to allocate orders and plan vehicle routes; wherein, the deep reinforcement learning algorithm transforms the temporal and spatial features of orders into feature vectors through an embedding layer, and inputs them into the spatial attention mechanism and temporal attention mechanism for processing respectively. After a preset number of iterations, a customer vector is generated through a gating fusion mechanism, and then the access order of each vehicle is determined through policy decoding to complete the order allocation and vehicle route planning for each route from the warehouse to the warehouse after completing the task.

[0042] In this embodiment, during the embedding layer processing, the start / end time of the time window and the service duration are normalized to a 0-1 numerical vector; the latitude and longitude of the node and its distance from the warehouse are converted into Euclidean distance or Manhattan distance vectors; the spatial attention mechanism extracts the dependency relationship of geographical proximity by calculating the similarity of spatial distances between nodes; the temporal attention mechanism extracts the dependency relationship of temporal constraints by calculating the overlap of time windows or the correlation of service durations; the gating fusion mechanism selects the nearest neighbor node according to the dynamic ratio of spatial and temporal weights, such as when the spatial weight accounts for 70%, and merges the two features to generate a customer vector; the policy decoder focuses on the customer vector of unassigned nodes through the attention mechanism based on the current state of the assigned tasks of the vehicle, outputs the probability of each node being selected, such as the node with the highest probability being the next access point, and finally generates a complete route starting from the warehouse, covering all task points and returning.

[0043] This embodiment captures the key relationships between task nodes in the spatial and temporal dimensions through a spatiotemporal attention mechanism, such as focusing on visiting neighboring nodes and prioritizing avoiding nodes with time conflicts. Then, it dynamically balances the influence of the two through gating fusion, and finally uses reinforcement learning to optimize and generate an efficient scheduling route through trial and error.

[0044] This embodiment considers both spatial and temporal factors, avoiding the problem of traditional methods that only focus on path length while ignoring time constraints, thus improving the rationality and efficiency of the route.

[0045] Step S104: Construct a reinforcement learning algorithm based on actor and critic, and train the target parameter network; wherein, the reinforcement learning algorithm initializes the actor policy network and the critic network, and through multiple rounds of training, calculates the comprehensive embedded features of customers in each round, generates vehicle routes, calculates the reward value of the route and the state value of the valuation network, updates the parameters of the actor policy network and the critic network based on this, completes the training for the preset number of rounds, and obtains the target parameter network.

[0046] In this embodiment, during initialization, the actor policy network and the critic network typically employ a multi-layer fully connected neural network structure. For example, the input layer receives customer embedding features, there are 2-3 hidden layers, and the output layers represent the probability distribution and state value, respectively. During training, in each round, customer embedding features are sampled from historical or simulation data and input into the actor network to generate task selection probabilities for each vehicle. The node with the highest probability is selected as the current action according to a greedy strategy. The action is then input into the simulation environment for execution, resulting in a new state, actual reward, and an end flag. The critic network calculates the state value based on the current state and the actual reward of the new state, and outputs the advantage function. The actor network updates its parameters based on the advantage function, and the critic network updates its parameters based on the error between the actual reward and the predicted value. This process is repeated until a preset number of training rounds is reached or the validation set performance stabilizes.

[0047] In this embodiment, the actor strategy network acts as the decision-maker, generating scheduling actions; the critic network acts as the evaluator, predicting the future benefits of the current state and providing actors with directions for improvement; gradually improving the rationality of the strategy. Through trial and error learning, it adapts to complex and ever-changing delivery scenarios; the critic's valuable feedback helps actors quickly identify inefficient strategies, accelerating convergence to the optimal solution; the final target parameter network can be directly deployed to a real-world system, achieving efficient real-time scheduling.

[0048] In one embodiment of the present invention, based on step S103, a possible embodiment will be given below, and its specific implementation will be described in a non-limiting manner.

[0049] Appendix Figure 3 This paper describes the overall structure of a deep reinforcement learning model based on a spatiotemporal attention mechanism. The model mainly consists of four core modules: an embedding layer, a spatiotemporal attention mechanism, a gating fusion mechanism, and a policy decoder. First, the model receives the raw input information of the task nodes and the vehicle state, and maps them into low-dimensional feature representations through the embedding layer. Based on this, the spatial attention module and the temporal attention module model the spatial distribution features and temporal window constraints of the task, respectively, and extract the dependencies between tasks through a multi-head attention mechanism.

[0050] Specifically, the spatial attention mechanism is calculated as follows: for any two task nodes i and j, their spatial correlation weight is given by the following formula:

[0051]

[0052]

[0053]

[0054]

[0055] The query vector represents the spatial multi-head attention, used to obtain the attention weights corresponding to that node. Used for calculation and Similarity or correlation between them Used to obtain the final output. It is a feedforward layer used to evolve node embeddings to obtain more information.

[0056] The time attention mechanism is calculated as follows: For any two task nodes i and j, their time correlation weights are given by the following formula:

[0057]

[0058]

[0059]

[0060]

[0061] The query vector represents the temporal multi-head attention, used to obtain the temporal attention weights between nodes. Used for calculation and Similarity or correlation between them This is used to obtain the final output. It is a feedforward layer used to evolve node embeddings to obtain more information.

[0062] Subsequently, a gating fusion mechanism is used to dynamically integrate the two types of features to generate a unified contextual representation. The gating fusion mechanism employs the following calculation method:

[0063]

[0064] During the policy decoding phase, the model uses local and global GRU modules to dynamically model the current vehicle state and generates the probability distribution of the currently available tasks through an attention-focusing mechanism. It guides the vehicle in selecting tasks and making path decisions in its current state.

[0065] like Figure 4 As shown, in step S103, the specific steps of deep reinforcement learning based on the spatiotemporal attention mechanism are as follows: Step 1: Initialize A0 = 0; Step 2: Enter the order collection ; Step 3: Extract the time feature information of the order. and spatial feature information ; Step 4. Input the spatial feature vector of the order into the spatial attention mechanism to generate the corresponding spatial representation vector. ; Step 5: Input the time feature vector of the order into the time attention mechanism to generate the corresponding time representation vector. ; Step 6: Check if the loop has run L times. If it has, jump to... Step 7. Otherwise, A1 = A0 + 1, jump to... Step 4; It's important to note that L represents the number of feature iterations in the spatiotemporal attention mechanism, used to control the optimization depth of the spatial and temporal attention modules for order features. In practical applications, L is often determined through experimental experience. For example, it can initially be set to 5-10 iterations. This is because 5 iterations avoid computational redundancy and improve efficiency, while 10 iterations ensure sufficient integration of spatial and temporal features, such as repeatedly reinforcing the correlation between geographical proximity and temporal constraints. During the training or debugging phase, the sufficiency of L can be judged by observing the changes in the feature vectors. If the feature vectors converge after L iterations, then the current L is a reasonable value.

[0066] Step 7: Gating fusion mechanism fuses temporal and spatial representation vectors to generate the final customer vector.

[0067] Step 8: Select vehicle k=0; Step 9: Generate context information for vehicle k ; Step 10: Generate context information for all vehicles ; Step 11: Generate the probability distribution between vehicle k and nodes ; Step 12: Greedily select the order with the highest probability as the next customer to visit vehicle k; Step 13: Check if all nodes have been allocated. If so, proceed. Step 15, otherwise execute Step 14; Step 14: Select the (k+1)%xth car and jump to... Step 10; Step 15: The algorithm ends, and the route plan is output.

[0068] The iterative design (l from 0 to L) in Steps 4-6 of this embodiment allows the spatial and temporal feature vectors to be continuously optimized through multiple rounds of attention calculation. In each iteration, spatial attention adjusts the spatial association weights based on the temporal features of the previous round. For example, when a node's time window is tight, its spatial weights with surrounding nodes are strengthened to shorten the path. Temporal attention also updates temporal dependencies based on the spatial features of the previous round. For example, when a node is close to the warehouse, its time window constraints are appropriately relaxed. This spatiotemporal feature feedback iterative mechanism upgrades feature representation from static shallow association to dynamic deep association, solving the problem that feature representation in existing technologies is difficult to adapt to complex scenarios.

[0069] The gating fusion mechanism in Step 7 can automatically adjust according to the real-time scenario: when the overall order time window is tight, z approaches 0; when the node spatial distribution is scattered, z approaches 1. This scenario-adaptive weight allocation capability enables customer vectors to accurately match the current scheduling requirements, overcoming the problem of poor feature adaptability caused by fixed weights.

[0070] Steps 8-14 drive allocation based on vehicle context information: Each vehicle's context includes its current load, elapsed time, and remaining capacity. When generating the probability distribution pi, orders matching the vehicle's status are prioritized, such as heavy cargo orders with sufficient remaining capacity and urgent orders whose time windows match the vehicle's elapsed time, assigning them a high probability. Simultaneously, a vehicle cyclical selection mechanism ensures balanced load across all vehicles. This vehicle status-order feature matching allocation logic solves the problem of order allocation being disconnected from actual vehicle capacity in existing technologies, improving resource utilization.

[0071] Step 12's greedy selection of the order with the highest probability ensures efficiency by prioritizing the use of the current optimal solution while reserving exploration space through dynamic updates of the probability distribution. For example, a low-probability order may become optimal in subsequent iterations due to the completion of other nodes. This strategy, which prioritizes utilization while supplementing exploration, avoids the trap of local optima and ensures the efficiency of route planning, solving the problem of balancing exploration and utilization in existing technologies.

[0072] The temporal and spatial features of orders collected in step S101, such as node latitude and longitude and region affiliation, are directly used as the raw input in Step 3. For example, spatial features include the distance and type between nodes in S101, enabling the spatial attention in Step 4 to calculate association weights based on this fundamental data; temporal features include the time window information of orders in S101, enabling the temporal attention in Step 5 to accurately capture temporal constraints. This fusion ensures that feature extraction is not divorced from actual data, providing a reliable foundation for subsequent attention calculations.

[0073] Constraint (2) automatically reduces the probability of the node in subsequent pi after selecting an order in Step 12; Constraint (8) assigns extremely low probability to orders exceeding the remaining capacity of the vehicle through the vehicle context passed to the probability distribution in Step 9; Constraint (4) automatically increases the association weight of the warehouse node in the probability distribution of the last order when pi is generated in Step 11.

[0074] The route planning results output in step S103, including vehicle access order, total distance, and timeout status, serve as samples for the experience replay pool in Step S104. The route reward value R(π) is calculated based on the actual performance in S103, and the customer vector from Step 7 serves as input to the comprehensive customer embedding features in S104. Simultaneously, the target parameter network trained in S104 feeds back into the generation of the probability distribution pi in S103, forming a closed loop of generation-training-optimization.

[0075] In one embodiment of the present invention, based on step S104, a possible embodiment will be given below, and its specific implementation will be described in a non-limiting manner.

[0076] like Figure 5 As shown, the specific steps of the reinforcement learning training algorithm based on actor critics in step S104 are as follows: Step 1: Initialize the actor strategy network and the network of critics ; Step 2: Set the training epoch=1; Step 3: Set batch=1; Step 4: Determine if Batch training cycles have been completed within an epoch. If so, proceed to... Step 14, otherwise jump to Step 5; Step 5: Calculate the customer's comprehensive embedded features through the encoder. ; Step 6: Generate vehicle routes using the decoder ; Step 7: Calculate the route Reward value ; Step 8: Calculate the state value of the valuation network ; Step 9. Using the sampled action sequences and rewards, estimate the gradient of the current policy network. ; Step 10: Calculate the squared advantage of the evaluation network ; Step 11: Update actor strategy network parameters ; Step 12: Update critic network parameters ; Step 13: Set batch = batch + 1, then jump to... Step 4; Step 14: Set epoch = epoch + 1; Step 15: Check if the epoch has been completed Epoch times. If so, jump to... Step 16, otherwise jump to Step 3; Step 16: The algorithm ends, and the actor policy network parameters are returned. .

[0077] In step S104 of this embodiment, the actor-critic algorithm simultaneously trains the policy network (actor θ) and the valuation network (critic): the actor network generates specific scheduling routes, and the critic network evaluates the future cumulative reward of the current state. This collaborative mechanism of policy generation and value evaluation enables the model to adjust the policy more accurately through the critic's feedback, avoiding the convergence difficulties caused by reward sparsity or delay in traditional single-network learning, and improving the optimization efficiency and stability of the scheduling policy.

[0078] The training process in step S104 (Step 1-Step 15) uses multiple epochs and batch training to expose the model to diverse delivery scenarios (such as different time periods and order densities). For example, in the early stages of training, the model may generate inefficient routes that are short but time-out. The critic network will use negative feedback to push the actor network to adjust its strategy. As training progresses, the model gradually learns to prioritize strategies that avoid timeouts when time windows are tight, and reinforces this pattern in subsequent batches. This dynamic learning mechanism of trial and error-feedback-optimization enables the policy network to adapt to unexpected situations in real delivery, improving generalization ability by more than 30% compared to traditional static scheduling methods.

[0079] The reward function design in step S104 (Step 7) is bound to the scheduling model constraints in step S102 (as shown in equations (1)-(11)): the reward value R(π) includes not only the reduction in path length (e.g., +5 minutes for every 1 km shortened), but also penalties for timeout costs (e.g., -10 minutes for every 1 hour of timeout) and penalties for exceeding vehicle load limits, directly corresponding to the multi-objective optimization requirements in the scheduling model. The state value assessment of the commentator network (Step 8) is based on these multi-objective constraints to predict the total revenue that the current route may obtain in the future. This deep binding of constraint → reward → value enables the model to actively avoid violations (e.g., overloading, timeout) during the optimization process, ensuring that the generated scheduling scheme conforms to actual business rules, and improving the compliance rate compared to the redundant process of planning first and then verifying in traditional methods.

[0080] The customer vector generated in step S103 serves as the input feature for step S104 (the comprehensive embedding features of customers in Step 5), directly participating in the action generation of the policy network (Step 11) and the value evaluation of the critic network (Step 8). For example, the temporal features in the customer vector (such as the start time of the time window) are used by the critic network to calculate the value loss due to timeout if the node is selected in the future, while the spatial features (such as the distance to the warehouse) are used by the actor network to generate the action probability of nearby delivery. This direct mapping between spatiotemporal features and policy value ensures that the route generated by the model not only conforms to the mathematical optimization objective but also to the geographical logic of actual delivery, avoiding the problem of mathematical optimization but practical infeasibility in traditional methods.

[0081] The scheduling model constraints in step S102 (such as the vehicle load limit in equations (7)-(8) and the time window constraint in equation (10)) are explicitly encoded into the reward function in step S104. For example, when the actor network generates a route where vehicle k exceeds the load limit, the reward function will directly deduct the corresponding score, and the critic network will predict the future loss caused by this violation when evaluating the state value. This intrinsic design of constraint → reward → value enables the model to learn the decision logic that compliance is a prerequisite and optimization is the goal during the training phase, without the need for additional constraint verification during the inference phase, which significantly improves the feasibility of the scheduling scheme.

[0082] Step S104's multi-round training (Step 15) is deeply integrated with the transportation network graph G=(V,E) constructed in Step S101: During training, the model will encounter scenarios corresponding to different edge weights in the network graph (such as the morning rush hour d). 0i (The value is relatively large, and d0i is relatively small at night), and through learning from the commentator network, it was found that d should be prioritized during the morning rush hour. 0i A strategy for shorter routes. This mapping between the physical meaning of the network graph and the contextualized strategy allows the routes generated by the model to be dynamically adjusted according to actual traffic conditions.

[0083] The actor network generates actions (routes), and the critic network evaluates the value of these actions (future rewards). These two networks guide each other through gradient updates (Step 11-Step 12), forming an exploration-exploitation balance and overcoming the local optima limitations of traditional single-network learning. The constraints of the scheduling model directly serve as the design basis for the reward function, making compliance a hard indicator for policy optimization rather than a post-hoc validation term, thus solving the pain point of traditional methods struggling to balance compliance and efficiency. The diversity of network graph edge weights and customer vectors encountered during training enables the model to learn optimal strategies for different time periods and regions, avoiding the limitations of traditional methods that rely on manual parameter adjustments based on experience.

[0084] In one embodiment of the present invention, based on step S103, the following will provide a possible embodiment and its specific implementation will be described in a non-limiting manner. Step S103 further includes: Step S1031: Extract multidimensional feature information of the order, including time features and spatial features, and transform these features into initial time feature vectors and initial spatial feature vectors respectively through an embedding layer, wherein the dimension of the feature vectors is dynamically adjusted according to the order size.

[0085] Optionally, temporal characteristics include the order's time window, deadline, and estimated processing time. Spatial characteristics include the node's latitude and longitude, region affiliation, and surrounding road density.

[0086] Step S1032: Input the initial spatial feature vector into the spatial attention mechanism to obtain the spatial correlation weights between different order nodes. The spatial correlation weights comprehensively consider the geographical distance of nodes, regional clustering, and spatial association patterns in historical transportation paths to generate a spatial representation vector containing the spatial dependencies of nodes. At the same time, input the initial temporal feature vector into the temporal attention mechanism, and calculate the temporal correlation weights by combining the urgency of the order's time window and the temporal dependency of the task type (pickup / delivery) (e.g., the pickup task must take priority over the corresponding delivery task) to generate a temporal representation vector containing temporal constraint associations.

[0087] Step S1033: Set the number of iterations L, which is adaptively adjusted according to the order complexity. In each iteration, update the spatial representation vector and the temporal representation vector respectively, so that the vectors gradually strengthen the spatiotemporal correlation features between nodes; after the iteration ends, the fusion weight of spatiotemporal features is dynamically adjusted through a gating fusion mechanism. The weights are adaptively allocated according to the spatial distribution density and the strictness of time constraints in the current task, generating a comprehensive customer vector containing multidimensional correlation features.

[0088] Step S1034: Combining the real-time status of each vehicle, such as current load, remaining capacity, and mileage traveled, with the customer comprehensive vector, the local GRU module captures the current task execution status of the vehicle, while the global GRU module controls the global optimality of the overall path planning, generating the access probability distribution between the vehicle and each node; based on the probability distribution and combined with vehicle load constraints and total route distance constraints, the access order of each vehicle is determined by a combination of greedy selection and backtracking adjustment, completing the order allocation, ensuring that each route starts from the warehouse, completes the tasks in sequence, and returns to the warehouse, and that the routes do not overlap or repeat.

[0089] As can be seen, by processing the spatiotemporal features of orders in stages, the attention mechanism is used to accurately capture the spatial clustering relationships and temporal dependencies between nodes, and the optimal path is generated by combining the real-time vehicle status. In terms of implementation, the temporal, spatial, and attribute features of the orders are first extracted and converted into vectors. The feature associations are strengthened through multiple rounds of spatiotemporal attention iteration. Then, the weights of spatiotemporal features are dynamically balanced according to the task scenario through a gating fusion mechanism. Finally, a route that meets the requirements of load, time, and distance is generated by combining the actual constraints of the vehicle. This solves the problems of spatiotemporal feature fragmentation and path planning ignoring the dynamic status of vehicles in traditional methods.

[0090] In one embodiment of the present invention, based on step S104, the following will provide a possible embodiment and its specific implementation will be described in a non-limiting manner. Step S104 specifically includes: Step S1041: Based on the input and output requirements of the express delivery vehicle scheduling, design the input layer, hidden layer, and output layer of the actor strategy network; simultaneously design the input layer, hidden layer, and output layer of the critic network. Network layers are connected through non-linear activation functions to ensure non-linear feature extraction capabilities.

[0091] Optionally, the input layer of the actor policy network can be the received customer embedding features. The hidden layers can contain multiple fully connected or recurrent neural network units. The output layer can output the probability distribution of task selection for each vehicle. The critic network's input layer can share customer embedding features with the actor network. The hidden layers are symmetrical to the actor network's hidden layer structure, and the output layer outputs the comprehensive value assessment of the current state.

[0092] Step S1042: Collect historical express order data, vehicle dispatch records and corresponding actual operation results, and generate simulated dispatch data in combination with the simulation environment; standardize the raw data to eliminate differences in units; divide the data into training set, validation set and test set, where the training set is used for model parameter learning, the validation set is used for adjusting hyperparameters, and the test set is used for final performance evaluation.

[0093] Step S1043: Define the reward function as a comprehensive evaluation index of the scheduling results, which includes positive rewards and negative penalties; the state value evaluation rule is the commentator network's prediction of the future cumulative reward of the state corresponding to the current customer's embedded features, and optimize the commentator network by comparing the deviation between the actual cumulative reward and the predicted value.

[0094] Step S1044: Initialize the weight parameters of the actor policy network and the critic network, such as using the Xavier initialization method; in each training round, randomly sample a batch of customer embedding features from the training set, input them into the actor network to generate the task selection probability distribution for each vehicle, and greedily select the task with the highest probability as the current action; input the action into the simulation environment or historical data simulator to obtain the new state, actual reward, and whether it has ended after execution; the critic network calculates the state value based on the actual reward of the current state and the new state, and outputs the advantage function, the difference between the actual reward and the predicted value; calculate the REINFORCE algorithm or PPO optimization method of the actor network based on the advantage function, and update the actor network parameters to maximize the cumulative reward; at the same time, use the mean square error between the actual reward and the critic's predicted value as the loss function to update the critic network parameters to improve the accuracy of value assessment; repeat the above process until the preset number of training rounds or the performance of the validation set converges, and finally save the optimal parameters as the target parameter network.

[0095] Step S104 of this embodiment uses an Actor-Critic reinforcement learning framework to jointly optimize the policy network and the valuation network, ultimately obtaining a target parameter network adaptable to complex scheduling scenarios. The implementation logic is as follows: the actor network, acting as the policy function, outputs the task selection probability of each vehicle based on the current time, space, and vehicle state, directly guiding scheduling decisions; the critic network, acting as the value function, evaluates the potential cumulative reward in the current state, providing feedback signals to the actor network to improve its strategy. During training, the two networks interact by sharing customer embedding features: the actor adjusts its strategy based on the difference between the critic's actual reward and the predicted value, increasing the probability of high-value actions; the critic updates the value prediction based on the actor's actual decision results, reducing the deviation between the predicted and actual values. Through multiple rounds of decision-feedback-optimization iterations, the actor network gradually learns the optimal scheduling strategy that minimizes the total transportation path and overtime cost under conditions such as time window constraints and vehicle capacity limitations, ultimately outputting a stable target parameter network that can be directly used for real-time scheduling of actual express delivery vehicles.

[0096] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0097] The following are embodiments of an intelligent scheduling system for logistics delivery vehicles based on a spatiotemporal attention mechanism provided in this disclosure. This system and the intelligent scheduling method for logistics delivery vehicles based on a spatiotemporal attention mechanism in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the intelligent scheduling system for logistics delivery vehicles based on a spatiotemporal attention mechanism, please refer to the embodiments of the intelligent scheduling method for logistics delivery vehicles based on a spatiotemporal attention mechanism described above.

[0098] The system includes: The data collection and network construction module is used to collect a set of express delivery orders, construct a transportation network diagram of express delivery vehicles, and obtain basic information about the set of transportation vehicles.

[0099] The scheduling model setting module is used to build a delivery vehicle scheduling model. The model objective is to minimize the sum of the total transportation route length and overtime cost of all vehicles, and constraints are set.

[0100] The path planning module is used to design a deep reinforcement learning algorithm based on a spatiotemporal attention mechanism for order allocation and vehicle route planning. The deep reinforcement learning algorithm transforms the temporal and spatial features of orders into feature vectors through an embedding layer, which are then processed by the spatial and temporal attention mechanisms respectively. After a preset number of iterations, a customer vector is generated through a gating fusion mechanism, and then the access order of each vehicle is determined through policy decoding. This completes the order allocation and vehicle route planning for each route, from starting from the warehouse to returning to the warehouse after completing the task.

[0101] The network training module is used to construct a reinforcement learning algorithm based on actor and critic networks and train the target parameter network. The reinforcement learning algorithm initializes the actor policy network and the critic network. Through multiple rounds of training, in each round, the comprehensive embedding features of the customer are calculated, vehicle routes are generated, the reward value of the route and the state value of the valuation network are calculated, and the parameters of the actor policy network and the critic network are updated based on this to complete the preset number of training rounds and obtain the target parameter network.

[0102] like Figure 6 As shown, this application also provides an electronic device, including a display module 103, a memory 102, a processor 101, and a computer program stored in the memory and executable on the processor 101. When the processor 101 executes the program, it implements the steps of an intelligent scheduling method for logistics delivery vehicles based on a spatiotemporal attention mechanism.

[0103] In embodiments of the present invention, electronic devices include, but are not limited to, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments described and / or claimed herein.

[0104] In this embodiment, processor 101 may be implemented using at least one of an application-specific integrated circuit, a programmable logic device, a field-programmable gate array, a processor, a controller, a microcontroller, a microprocessor, or an electronic unit designed to perform the functions described herein. In some cases, such an implementation may be implemented within a controller. For software implementation, implementations such as processes or functions may be implemented with separate software modules that allow the performance of at least one function or operation. Software code may be implemented by a software application (or program) written in any suitable programming language, and the software code may be stored in memory and executed by the controller.

[0105] The display module 103 is used to display information input by the user or information provided to the user. The display module 103 may include a display panel, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like.

[0106] The memory 102 can be used to store software programs and various data. The memory 102 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0107] This application also provides a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the intelligent scheduling method for logistics delivery vehicles based on a spatiotemporal attention mechanism.

[0108] The storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example,, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0109] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for intelligent scheduling of logistics delivery vehicles based on a spatiotemporal attention mechanism, characterized in that, The methods include: S101: Collect a set of express delivery orders, construct a transportation network diagram of express delivery vehicles, and obtain basic information about the set of transportation vehicles; S102: Construct a delivery vehicle scheduling model, setting the model objective as minimizing the sum of the total transportation route length and overtime cost of all vehicles, and setting constraints. S103: Design a deep reinforcement learning algorithm based on spatiotemporal attention mechanism for order allocation and vehicle route planning; wherein, the deep reinforcement learning algorithm transforms the temporal and spatial features of the order into feature vectors through an embedding layer, and inputs them into the spatial attention mechanism and temporal attention mechanism for processing respectively. After a preset number of iterations, a customer vector is generated through a gating fusion mechanism, and then the access order of each vehicle is determined through policy decoding to complete order allocation and vehicle route planning for each route from the warehouse to the warehouse after completing the task; S104: Construct a reinforcement learning algorithm based on actor and critic networks to train the target parameter network; wherein, the reinforcement learning algorithm initializes the actor policy network and the critic network, and through multiple rounds of training, calculates the comprehensive embedding features of customers in each round, generates vehicle routes, calculates the reward value of the routes and the state value of the valuation network, updates the parameters of the actor policy network and the critic network based on this, completes the training for the preset number of rounds, and obtains the target parameter network.

2. The intelligent scheduling method for logistics delivery vehicles based on spatiotemporal attention mechanism according to claim 1, characterized in that, Step S101 specifically includes: Constructing a transportation network map for express delivery vehicles , where V contains Warehouse node, delivery node and pickup nodes , This represents the set of delivery nodes and pickup nodes. For the set of edges between nodes, each edge This represents the optimal route from node i to node j. Let represent the distance to edge (i,j), where the distance satisfies symmetry, i.e. The distance between the same nodes is 0, that is , ; Let x represent the set of available vehicles, and x be the number of vehicles. This indicates the maximum load capacity of vehicle k.

3. The intelligent scheduling method for logistics delivery vehicles based on spatiotemporal attention mechanism according to claim 1, characterized in that, Defined in step S102 Indicates vehicle From point Drive directly to the destination If so, then ,otherwise ; The established express delivery vehicle dispatching model is as follows: (1) The constraints are set as follows: (2) (3) (4) (5) (6) Equation (1) is the optimization objective function, which represents minimizing the total transportation routes and overtime costs for all vehicles; Equation (2) indicates that each client node is accessed only once; Equation (3) represents the vehicle k Visit a customer; Equation (4) indicates that the vehicle will definitely depart from the warehouse and return to the warehouse; Equations (5) and (6) indicate that after meeting customer needs, the vehicle... k The capacity is adjusted accordingly; Equation (7) indicates that vehicle k will carry the requirements of all delivery nodes when it departs from the warehouse; Equation (8) indicates that the vehicle capacity cannot be exceeded by the total demand on a single route; Equation (9) indicates that if the vehicle k From node i To the node j The arrival time must be greater than the node. i Arrival time plus node i The service time of a node is expressed by Equation (10), which indicates that the service time of a node is no earlier than the earliest service time window of the node. In the formula, M is a positive number, and β is the weighting coefficient for timeout cost; The time it takes for vehicle k to arrive at node j; The time window cutoff time for node j; i and j Represents any two nodes, used to describe the relationship between the nodes; Let k be the time it takes for vehicle k to arrive at node i. Let be the travel time of vehicle k from node i to node j; The remaining load capacity of vehicle k after the operation at node j; Let K be the remaining load capacity of vehicle k after the operation at node i. Let K be the initial load capacity of vehicle k when it departs from the warehouse. Let be the demand for goods at node i; The service time for node i; Let k be the service time from node i to node j. This represents the earliest service time window for node i.

4. The intelligent scheduling method for logistics delivery vehicles based on spatiotemporal attention mechanism according to claim 1, characterized in that, In step S103, the specific steps of deep reinforcement learning based on the spatiotemporal attention mechanism are as follows: Step 1: Initialize A0 = 0; Step 2: Enter the order collection ; Step 3: Extract the time feature information of the order. and spatial feature information ; Step 4. Input the spatial feature vector of the order into the spatial attention mechanism to generate the corresponding spatial representation vector. ; Step 5: Input the time feature vector of the order into the time attention mechanism to generate the corresponding time representation vector. ; Step 6: Check if the loop has run L times. If it has, jump to... Step 7. Otherwise, A1 = A0 + 1, jump to... Step 4; Step 7: Gating fusion mechanism fuses temporal and spatial representation vectors to generate the final customer vector. Step 8: Select vehicle k=0; Step 9: Generate context information for vehicle k ; Step 10: Generate context information for all vehicles ; Step 11: Generate the probability distribution between vehicle k and nodes ; Step 12: Greedily select the order with the highest probability as the next customer to visit vehicle k; Step 13: Check if all nodes have been allocated. If so, proceed. Step 15, otherwise execute Step 14; Step 14: Select the (k+1)%xth car and jump to... Step 10; Step 15: The algorithm ends, and the route plan is output.

5. The intelligent scheduling method for logistics delivery vehicles based on spatiotemporal attention mechanism according to claim 1, characterized in that, The specific steps of the reinforcement learning training algorithm based on actor critics in step S104 are as follows: Step 1: Initialize the actor strategy network and the network of critics ; Step 2: Set the training epoch=1; Step 3: Set batch=1; Step 4: Determine if Batch training cycles have been completed within an epoch. If so, proceed to... Step 14, otherwise jump to Step 5; Step 5: Calculate the customer's comprehensive embedded features through the encoder. ; Step 6: Generate vehicle routes using the decoder ; Step 7: Calculate the route Reward value ; Step 8: Calculate the state value of the valuation network ; Step 9. Using the sampled action sequences and rewards, estimate the gradient of the current policy network. ; Step 10: Calculate the squared advantage of the evaluation network ; Step 11: Update actor strategy network parameters ; Step 12: Update critic network parameters ; Step 13: Set batch = batch + 1, then jump to... Step 4; Step 14: Set epoch = epoch + 1; Step 15: Check if the epoch has been completed Epoch times. If so, jump to... Step 16, otherwise jump to Step 3; Step 16: The algorithm ends, and the actor policy network parameters are returned. .

6. The intelligent scheduling method for logistics delivery vehicles based on spatiotemporal attention mechanism according to claim 1, characterized in that, Step S103 also includes: Extract multidimensional feature information from orders, and transform the multidimensional feature information into initial temporal feature vector and initial spatial feature vector through an embedding layer; The initial spatial feature vector is input into the spatial attention mechanism to obtain the spatial correlation weights between different order nodes. The spatial correlation weights comprehensively consider the geographical distance of nodes, regional clustering, and spatial association patterns in historical transportation paths to generate a spatial representation vector containing the spatial dependencies of nodes. The initial time feature vector is input into the time attention mechanism, and the time correlation weight is calculated by combining the time window urgency of the order and the temporal dependency of the task type to generate a time representation vector containing time constraint associations. Set the number of iterations L, and update the spatial representation vector and temporal representation vector in each iteration to gradually strengthen the spatiotemporal correlation features between nodes; after the iteration ends, the fusion weight of spatiotemporal features is dynamically adjusted through a gating fusion mechanism to generate a comprehensive customer vector containing multidimensional correlation features. By combining the real-time status of each vehicle with the customer's comprehensive vector, the local GRU module captures the current task execution status of the vehicle and generates the access probability distribution between the vehicle and each node. Based on the probability distribution and combined with vehicle load constraints and total route distance constraints, the access order of each vehicle is determined by a combination of greedy selection and backtracking adjustment to complete the order allocation.

7. The intelligent scheduling method for logistics delivery vehicles based on spatiotemporal attention mechanism according to claim 1, characterized in that, Step S104 specifically includes: Based on the input and output requirements of express delivery vehicle scheduling, the input layer, hidden layer, and output layer of the actor strategy network are designed; the input layer, hidden layer, and output layer of the critic network are designed simultaneously; the network layers are connected by nonlinear activation functions. Collect historical express delivery order data, vehicle dispatch records, and corresponding actual operation results, and generate simulated dispatch data in combination with the simulation environment; standardize the raw data; divide the data into training set, validation set, and test set; The reward function is defined as a comprehensive evaluation index of the scheduling results, including positive rewards and negative penalties; the state value evaluation rule is the commentator network's prediction of the future cumulative reward of the state corresponding to the current customer's embedded features, and the commentator network is optimized by comparing the deviation between the actual cumulative reward and the predicted value. Initialize the weight parameters of the actor policy network and the critic network. In each training round, randomly sample a batch of customer embedding features from the training set and input them into the actor network to generate the task selection probability distribution for each vehicle. Greedily select the task with the highest probability as the current action. Input the action into the simulation environment or historical data simulator to obtain the new state, actual reward, and termination flag after execution. The critic network calculates the state value based on the current state and the actual reward of the new state, and outputs the advantage function. Calculate the policy gradient of the actor network based on the advantage function and update the actor network parameters to maximize the cumulative reward. At the same time, use the mean squared error between the actual reward and the critic's predicted value as the loss function to update the critic network parameters. Repeat the above process until the preset number of training rounds or validation set performance convergence is reached. Finally, save the optimal parameters as the target parameter network.

8. A smart dispatching system for logistics delivery vehicles based on a spatiotemporal attention mechanism, characterized in that, The system is used to implement the intelligent scheduling method for logistics delivery vehicles based on spatiotemporal attention mechanism as described in any one of claims 1 to 7; The system includes: The data collection and network construction module is used to collect a set of express delivery orders, construct a transportation network diagram of express delivery vehicles, and obtain basic information about the set of transportation vehicles. The scheduling model setting module is used to build a delivery vehicle scheduling model. The model objective is to minimize the sum of the total transportation route length and overtime cost of all vehicles, and constraints are set. The path planning module is used to design a deep reinforcement learning algorithm based on a spatiotemporal attention mechanism for order allocation and vehicle route planning. The deep reinforcement learning algorithm transforms the temporal and spatial features of orders into feature vectors through an embedding layer, which are then processed by the spatial and temporal attention mechanisms respectively. After a preset number of iterations, a customer vector is generated through a gating fusion mechanism, and then the access order of each vehicle is determined through policy decoding. This completes the order allocation and vehicle route planning for each route, from starting from the warehouse to returning to the warehouse after completing the task. The network training module is used to construct a reinforcement learning algorithm based on actor and critic networks and train the target parameter network. The reinforcement learning algorithm initializes the actor policy network and the critic network. Through multiple rounds of training, in each round, the comprehensive embedding features of the customer are calculated, vehicle routes are generated, the reward value of the route and the state value of the valuation network are calculated, and the parameters of the actor policy network and the critic network are updated based on this to complete the preset number of training rounds and obtain the target parameter network.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the intelligent scheduling method for logistics delivery vehicles based on the spatiotemporal attention mechanism as described in any one of claims 1 to 7.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the intelligent scheduling method for logistics delivery vehicles based on the spatiotemporal attention mechanism as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Vehicle scheduling method and device

    CN113449895A

  • Take-out delivery path planning method based on deep reinforcement learning

    CN115841286A

  • Malicious traffic detection method based on integral space-time diagram convolutional neural network fused with space-time attention

    CN117579290A

  • Logistics order management optimization method and device, equipment and storage medium

    CN120218782A

  • Traffic signal cooperative control method based on multi-agent reinforcement learning

    CN120340272A

Cited By

  • Dynamic vehicle path planning method and system, electronic equipment and storage medium

    CN121581356A

  • Method and apparatus for constructing vehicle routes with time windows

    CN122414958A