Learning type electric vehicle goods taking and delivering scheduling method based on graph structure perception

By adopting a learning method based on graph structure perception in the electric vehicle pick-and-delivery path planning, combining Markov decision-making process and Reinforce reinforcement learning algorithm, the problems of low computing efficiency and high model complexity in the existing technology are solved, and the optimal path planning with the smallest energy consumption is achieved, which improves vehicle energy saving and logistics efficiency.

CN120124830APending Publication Date: 2025-06-10SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510203714.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art has problems of low computing efficiency and high model complexity in the planning of electric vehicle pick-up and delivery paths, especially in terms of dynamic demand and energy consumption optimization.

Method used

Using a learning method based on graph structure perception, a neural network model is constructed to optimize the delivery path of electric vehicles by constructing a map model of dynamic requirements, combining Markov decision-making process and Reinforce reinforcement learning algorithm. This method introduces a graph attention mechanism and feature embedding module, which improves the computing efficiency and the accuracy of path planning.

Benefits of technology

It is realized that the optimal path to find the smallest energy consumption for the electric fleet to complete the pick-up and delivery task while ensuring customer time and vehicle battery power, and improves the vehicle's energy saving level and logistics and transportation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124830A_ABST
    Figure CN120124830A_ABST
Patent Text Reader

Abstract

The invention provides a learning type electric vehicle goods taking and delivery scheduling method based on graph structure perception. The method comprises the following steps: establishing an electric vehicle goods taking and delivery scheduling traffic map model; constructing an accurate electric vehicle energy consumption model considering terrain and vehicle factors; describing a routing decision by adopting a Markov decision process based on the model; constructing a neural network model of an electric vehicle goods taking and delivering problem; and training the model by adopting a deep reinforcement learning algorithm. According to the method, aiming at the vehicle routing problem, the routing problem of the two stages of goods taking and goods delivery is solved in combination with the reality, high real-time performance is achieved for the electric vehicle routing decision of the delivery and pickup problems such as city logistics, take-out and taxi taking in practical application, and the dispatching efficiency of the goods taking and delivery problems of the electric vehicle is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle scheduling management. Specifically, it is a learning-based electric vehicle pickup and delivery scheduling method based on graph structure perception, which finds the optimal path with the minimum energy consumption for an electric vehicle fleet to complete pickup and delivery tasks under the conditions of ensuring customer time and vehicle battery power. Background Technique

[0002] Green logistics is currently the best development trend in the logistics industry cost, and electric vehicles are the key development objects of green logistics. Electric vehicles have the advantages of environmental protection and zero pollution, but they are also restricted by technical factors such as limited driving range and vehicle charging time. At the same time, customer needs are dynamically changing. Therefore, it is very important to optimize the management of electric vehicles to achieve lower energy consumption under technical constraints. The path planning of electric vehicles with the goal of minimum energy consumption is an effective means to reduce costs and increase efficiency for electric vehicle pickup and delivery.

[0003] In order to be more in line with the actual situation, the influence of factors such as terrain, acceleration, and braking is introduced into the energy consumption of electric vehicles, an accurate energy consumption model is established, road influence and charging station distribution are considered, and this energy consumption is incorporated into the electric vehicle pickup and delivery path problem. Developing a high-precision integrated electric vehicle economic routing method with dynamic demand is crucial for planning the best route for electric vehicle pickup and delivery. Existing learning-based methods face a large number of repeated calculations in the iterative operation of a large amount of data during the process of solving the optimal path, which also increases the complexity of the model. Therefore, improving the structure of the model to improve its calculation efficiency is also an important link in improving this technology. Summary of the Invention

[0004] Aiming at the deficiencies of the existing technology, the present invention provides a learning-based electric vehicle pickup and delivery scheduling method based on graph structure perception, aiming to improve the vehicle energy-saving level and logistics transportation efficiency, and reduce its complexity while improving the calculation efficiency of the controller.

[0005] Technical Solution: To solve the above problems, the present invention provides the following technical solution:

[0006] The present invention discloses a learning-based electric vehicle pickup and delivery scheduling method based on graph structure perception, including the following steps:

[0007] Step 1: Construct a map model G=(V, E) with dynamic demand. G is a complete directed graph including nodes and edges between nodes, where V represents the vertex set, including the parking lot set D, the customer set C (including two types of customer points: pickup and delivery), and the charging station set F. E is the set of edges in the graph; at the same time, set the vehicle path problem constraint conditions and objective function, and build a mathematical model for the electric vehicle pickup and delivery scheduling problem;

[0008] Step 2: Set the objective function as the energy consumption factor. Comprehensively consider the details of the energy consumption of electric vehicles, such as acceleration, speed, terrain, and the efficiency of the power system, and embed the above path details into the energy cost coefficient linearly related to the mass to establish an energy consumption model E between any two points i and j of the vehicle. ij ;

[0009] Step 3: Construct a Markov decision process for vehicle routing decisions;

[0010] Step 4: Construct a neural network model for the vehicle pick-up and delivery path problem according to the decision process;

[0011] The policy network consists of an encoder and a decoder. In the encoder, we innovatively add a graph attention mechanism, and we also add an adjacency matrix and a feature embedding module to process node features; this network processes the input current state and outputs the selection probability of the node;

[0012] The baseline network uses a structure similar to the policy network, selects the action with the highest probability using the greedy policy, reduces the variance in the policy network, and improves the stability and efficiency of training;

[0013] Step 5: Design a Reinforce reinforcement learning algorithm to train the network model in Step 4.

[0014] As a further preferred solution, in Step 1, the constructed mathematical model of the problem is:

[0015]

[0016] τ 0 = 0 (7)

[0017]

[0018]

[0019] u 0 = 0 (10)

[0020]

[0021] where x ij is a binary decision variable. When the vehicle departs from node i and visits node j, x ij = 1, otherwise 0; q j represents the demand of customer point j; d ij represents the Euclidean distance between node i and node j; v ij represents the average speed of the vehicle between node i and node j; t ij represents the travel time of the vehicle between node i and node j; E ijDenote the energy consumption of the vehicle between node i and node j; u ij Denote the load of the vehicle between node i and node j; U represents the maximum capacity of the vehicle; Q represents the battery capacity of the vehicle; g j Denote the service time when the vehicle arrives at node j; e j Denote the remaining battery power after the vehicle arrives at node j; τ j Denote the time required for the vehicle to reach node j;

[0022] The objective function (1) of the model represents that under the constraints of (2) to (18), the driving cost of the electric vehicle is minimized; among them, constraint (2) means that each customer is visited only once; constraint (3) means that not all charging stations need to be visited; constraint (4) represents the flow conservation; constraint (5) means that the number of available vehicles does not exceed K; constraint (6) means sorting according to the vehicle access vertex time; constraints (7) and (8) mean that the electric vehicle starts from the parking lot at the beginning, and the single-vehicle driving time cannot exceed the maximum time constraint; constraint (9) represents the real-time cargo load of the electric vehicle; constraints (10) and (11) mean that the electric vehicle is in an empty state when leaving the garage and the vehicle load cannot exceed the maximum load; constraints (12) and (13) mean tracking the vehicle's SOC state and ensuring that the battery power is within the normal range; constraint (14) means that the vehicle leaves the garage with a full charge and is fully charged when leaving the charging station; constraint (15) represents the decision variable constraint; constraint (16) represents the calculation of the arrival time of two adjacent nodes in the path; constraint (17) is the priority constraint to ensure that the arrival time of the pickup point is earlier than the arrival time of the delivery node; finally, constraint (18) imposes the non-negativity of the arrival time.

[0023] As a further preferred solution, in step 2, comprehensively considering details such as the acceleration, speed, terrain, and power system efficiency of the vehicle on all sections of the road, first determine the mechanical power P of the vehicle on the driving section ij , which is calculated as follows:

[0024] P ij =(m ij ·a + m ij ·g·(sin(α ij ) + C r ·cos(α ij ) + 0.5·C d ·ρ·A·v ij 2 )·v ij

[0025] In the formula, m ij represents the total mass of the vehicle driving on the edge (i, j), a represents the acceleration of the vehicle, A represents the frontal area of the EV, C dis the aerodynamic drag coefficient, C r is the rolling resistance coefficient, ρ is the air density, g is the gravitational constant, α ij is the average road slope between vertex i and vertex j, v ij represents the average speed between vertex i and vertex j;

[0026] Use t ij to represent the driving time of the vehicle along (i,j), φ d and φ r respectively represent the power of the motor during propulsion and regenerative braking, and respectively represent the charge and discharge efficiency of the battery, and the vehicle energy consumption E ij is calculated as follows:

[0027]

[0028] Based on the above formula, the energy consumption cost of all road segments in the directed graph in step 1 can be calculated.

[0029] As a further preferred solution, in step 3, the mathematical model of the electric vehicle pick-up and delivery scheduling problem is converted into a Markov decision process;

[0030] Define the vehicle state S t =(x t ,v t ) represents that at time step t, the system state consists of the graph vertex state and the vehicle state v t =[e t ,τ t ,u t ; u t ,e t ,τ t respectively represent static information, dynamic information, vertex two-dimensional coordinates, current vertex demand, load, power state, driving time;

[0031] Determine the access to the node that the vehicle should select in the current state, and the feasible action a t , and finally generate the action sequence {a 0 ,a 1 ,…,a T}.

[0032] Considering the dynamic demand, establish the state transition equation from node i to node j, and the next state S t+1 is updated according to the current state S t , and the vehicle state is updated after executing the action a t .

[0033] The reward function is expressed as where r t is expressed as the negative value of the incremental travel energy consumption at a step t.

[0034] As a further preferred solution, in step 4, the neural network model is composed as follows:

[0035] The policy network consists of an encoder and a decoder, and we add an adjacency matrix and a feature embedding module to process node features. After giving the graph information, the parameterized policy is expressed as:

[0036]

[0037] The graph attention mechanism learns the graph structure information in the encoder, transmits the graph information through the adjacency matrix, reduces the feature calculation amount, and trains a weight matrix W ∈ R for each node after giving the input F×F′ , then passes through LeakyReLU to provide non-linearity and uses the softmax function for normalization, calculates the attention coefficient of the graph attention layer for the node pair (i, j), and then updates its own feature as the input by calculating the linear combination of the expected adjacent neighbor node features. The following is the calculation equation system:

[0038]

[0039] To stabilize the learning process, a multi-head attention mechanism is also added, and an averaging operation is performed on the output using a non-linear function, as shown in the following formula:

[0040]

[0041] The adjacency matrix and the feature embedding module jointly process the node feature information, screen out the reachable valid nodes through the adjacency matrix, and the feature embedding module will perform linear projection and max pooling operations on the vehicle state S t and the routing feature R t , and perform feature fusion to obtain the output feature embedding O t , and then the feature embedding step will concatenate the vehicle state feature and the routing feature and pass them to the decoder;

[0042] The decoder generates a probability vector to select vertices, and through the context vector composed of the graph embedding, the embedding at time step t - 1, and the feature embedding O t and the multi-head attention calculation outputs a probability vector at each time step to represent the probability of each node being selected. The specific calculation formula group is shown as follows and outputs a probability vector at each time step to represent the probability of each node being selected. The specific calculation formula group is shown as follows

[0043]

[0044] Calculating a glimpse of context information

[0045]

[0046] where are all trainable matrices. Given that and k t = W K h N , so the compatibility calculation of the single-head attention h g with all nodes in step t is calculated as follows:

[0047]

[0048] Finally, we use the softmax function to calculate the probability vector and output the selection probability of the vertex, which is expressed as follows:

[0049] p t = softmax(h t + Z·Mask t )

[0050] In the formula, Mask t represents the masking rule, which is 1 if the vertex is masked and 0 otherwise;

[0051] The baseline network, using a structure similar to the policy network, selects the action with the highest probability using the greedy policy, reduces the variance in the policy network, and improves the stability and efficiency of training.

[0052] As a further preferred solution, in step 5, the training process of the Reinforce reinforcement learning algorithm is as follows:

[0053] The Reinforce reinforcement learning algorithm is used to train the model. The training objective is to find the policy parameter θ that maximizes the reward. The objective training function can be expressed as:

[0054]

[0055] In the formula, R(π θ ) represents the cumulative total reward, p θ (a|s) represents the probability of selecting a node, represents the logarithmic gradient of the probability of selecting an action.

[0056] As a further preferred solution, the examples during training are all generated data. In each time step, the Actor network generates probabilities and updates the state until the end; during the learning process, the model updates the network parameters according to the gradient; the graph attention mechanism is used to focus on the graph structure information to reduce the feature calculation of invalid points, and the feature embedding module is used to perform feature fusion in advance to improve the calculation efficiency; the Reinforce reinforcement learning algorithm is selected to retain the Actor network and combine it with the Baseline for policy optimization, reducing the variance in the Actor network and improving the stability and efficiency of training.

[0057] Beneficial effects

[0058] Compared with the prior art, the present invention has the following advantages:

[0059] 1. Based on the terrain environment of the road and the dynamic load information of the vehicle, the present invention constructs an accurate energy consumption model for electric vehicles, making it more in line with the actual situation.

[0060] 2. The present invention converts the pick-up and delivery problem into a Markov decision process, which is convenient for the Reinforce algorithm to solve, making the dynamic electric vehicle pick-up and delivery problem model universal for the reinforcement learning algorithm.

[0061] 3. Based on the encoder-decoder structure, the present invention constitutes a neural network architecture, introduces the graph attention mechanism, and adds the adjacency matrix and the feature embedding module, which can more effectively learn the graph structure information, has real-time response to dynamic information, improves the solution accuracy and calculation efficiency, and improves the overall energy efficiency.

[0062] 4. The present invention learns the graph structure information through the graph attention mechanism, and the proposed method is applicable to similar path solving problems. Brief description of the drawings

[0063] Figure 1 is the model of the electric vehicle pick-up and delivery path problem in the embodiment of the present invention;

[0064] Figure 2 is the flowchart of the scheduling method in the embodiment of the present invention;

[0065] Figure 3 is the framework diagram of the Reinforce algorithm in the embodiment of the present invention. Detailed implementation manners

[0066] The present invention will be further described below with reference to the accompanying drawings:

[0067] This embodiment provides a learning-based electric vehicle pick-up and delivery scheduling method based on graph structure perception. By constructing an electric vehicle energy consumption model, a Markov decision process for the electric vehicle pick-up and delivery path problem is established. Then, a neural network model based on the graph attention mechanism is innovatively constructed. Finally, an economic routing scheme is solved through the Reinforce reinforcement learning algorithm.

[0068] For the vehicle pick-up and delivery path problem model, refer to Figure 1 .

[0069] For the scheduling method, refer to Figure 2 , which specifically includes the following steps;

[0070] In step 1, the established mathematical model of the problem is:

[0071]

[0072]

[0073] τ 0 = 0 (7)

[0074]

[0075] u 0 = 0 (10)

[0076]

[0077] where x ij is a binary decision variable. When the vehicle departs from node i and visits node j, x ij = 1, otherwise 0; q j represents the demand of customer point j; d ij represents the Euclidean distance between node i and node j; v ij represents the average speed of the vehicle between node i and node j; t ij represents the driving time of the vehicle between node i and node j; E ij represents the energy consumption of the vehicle between node i and node j; u ij represents the load of the vehicle between node i and node j; U represents the maximum capacity of the vehicle; Q represents the battery capacity of the vehicle; g j represents the service time of the vehicle arriving at node j; e j represents the remaining battery power after the vehicle arrives at node j; τ j represents the time required for the vehicle to reach node j.

[0078] The objective function (1) of the model represents that the driving cost of the electric vehicle is minimized under the constraints (2) to (18); among them, constraint (2) means that each customer is visited only once; constraint (3) means that not all charging stations need to be visited; constraint (4) represents flow conservation; constraint (5) means that the number of available vehicles does not exceed K; constraint (6) means sorting according to the vehicle access vertex time; constraints (7) and (8) mean that the electric vehicle starts from the parking lot at the beginning, and the single-vehicle driving time cannot exceed the maximum time constraint; constraint (9) represents the real-time cargo load of the electric vehicle; constraints (10) and (11) mean that the electric vehicle is unloaded when leaving the garage and the vehicle load cannot exceed the maximum load; constraints (12) and (13) mean tracking the vehicle's SOC state and ensuring that the battery power is within the normal range; constraint (14) means that the vehicle leaves the garage with a full charge and is fully charged when leaving the charging station; constraint (15) represents the decision variable constraint; constraint (16) represents the calculation of the arrival time of two adjacent nodes in the path; constraint (17) is a precedence constraint to ensure that the arrival time of the pickup point is earlier than the arrival time of the delivery node; finally, constraint (18) imposes the non-negativity of the arrival time.

[0079] In step 2, considering in detail the acceleration, speed, terrain, and power system efficiency of the vehicle on all sections, first determine the mechanical power P of the vehicle on the driving section ij , which is calculated as follows:

[0080] P ij = (m ij · a + m ij · g · (sin(α ij ) + C r · cos(α ij ) + 0.5 · C d · ρ · A · v ij 2 ) · v ij

[0081] In the formula, m ij represents the total mass of the vehicle driving on the edge (i, j). a represents the acceleration of the vehicle. A represents the frontal area of the EV, C d is the aerodynamic drag coefficient, C r is the rolling resistance coefficient, ρ is the air density, g is the gravitational constant, α ij is the average road slope between vertex i and vertex j, and v ij represents the average speed between vertex i and vertex j.

[0082] Use t ij to represent the driving time of the vehicle along (i, j), φ d and φ rrespectively represent the power of the motor during propulsion and regenerative braking, and respectively represent the charge and discharge efficiency of the battery, and the vehicle energy consumption E ij is calculated as follows:

[0083]

[0084] Based on the above formula, the energy consumption cost of all road sections in the directed graph in step 1 can be calculated.

[0085] In step 3, the model is converted into a Markov decision process.

[0086] Define the vehicle state S t =(x t , v t ) represents that at time step t, the system state consists of the graph vertex state and the vehicle state v t =[e t , τ t , u t . u t , e t , τ t respectively represent static information, dynamic information, two-dimensional coordinates of the vertex, current vertex demand, load, power status, and driving time;

[0087] Determine the access to the node that the vehicle needs to select in the current state, and the feasible action a t , and finally generate the action sequence {a 0 , a 1 , …, a T}.

[0088] Considering the dynamic demand, establish the state transition equation from node i to node j. The next state S t+1 is updated according to the current state S t , and the vehicle state is updated after executing the action a t .

[0089] The reward function is expressed as where r t represents the negative value of the incremental travel energy consumption at a step t.

[0090] In step 4, the neural network model is composed as follows:

[0091] The policy network consists of an encoder and a decoder, and we add an adjacency matrix and a feature embedding module to process node features. After given the graph information, the parameterized policy is expressed as:

[0092]

[0093] The graph attention mechanism learns graph structure information in the encoder, transmits graph information through the adjacency matrix, reduces the amount of feature calculation, and trains a weight matrix W∈R for each node after the given input F×F′ , and then provides non-linearity through LeakyReLU and normalizes it using the softmax function to calculate the attention coefficient of the graph attention layer for the node pair (i,j), and then updates its own features as the input by calculating the linear combination of the features of the expected adjacent neighboring nodes. The following are the calculation equations:

[0094]

[0095] To stabilize the learning process, a multi-head attention mechanism is also added, and a non-linear function is used to average the output, as shown in the following formula:

[0096]

[0097] The adjacency matrix and the feature embedding module jointly process the node feature information. The effective nodes that can be reached are screened out through the adjacency matrix. The feature embedding module will perform linear projection and max pooling operations on the vehicle state S t and the routing feature R t , and after feature fusion, the output feature embedding O t is obtained. Then the feature embedding step will concatenate the vehicle state feature and the routing feature and pass them to the decoder.

[0098] The decoder generates a probability vector to select vertices, and through the context vector composed of the graph embedding, the embedding at time step t-1, and the feature embedding O t and the multi-head attention calculation outputs a probability vector at each time step to represent the probability of each node being selected. The specific calculation formula group is as follows Calculate the glimpse of the context information

[0099]

[0100]

[0101]

[0102] where are all trainable matrices; given that and k t =W K h N , so the compatibility calculation of the single-head attention h g with all nodes in step t is as follows:

[0103] ​

[0104] Finally, we use the softmax function to calculate the probability vector and output the selection probability of the vertex, which is expressed as follows:

[0105] p t = softmax(h t + Z·Mask t )

[0106] where Mask t represents the masking rule. If the vertex is masked, it is represented as 1; otherwise, it is 0.

[0107] The baseline network uses a structure similar to the policy network and selects the action with the highest probability using the greedy policy to reduce the variance in the policy network and improve the stability and efficiency of training.

[0108] In step 5, the training process of the Reinforce reinforcement learning algorithm is as follows:

[0109] The model is trained using the Reinforce reinforcement learning algorithm. The training objective is to find the policy parameter θ that maximizes the reward, and the objective training function can be expressed as:

[0110]

[0111] where R(π θ ) represents the cumulative total reward, p θ (a|s) represents the probability of selecting a node, represents the logarithmic gradient of the probability of selecting an action.

[0112] In the description of this specification, references to terms such as "one embodiment", "example", "specific example", etc. refer to those described in connection with that embodiment or example. In this specification, the above terms do not necessarily refer to the same embodiment or example. And the described features can be combined in a suitable manner in one or more embodiments or examples.

[0113] The above are the basic principle features and advantages of the present invention. The present invention is not limited by the above embodiments.

Claims

1. A learning-based electric vehicle pickup and delivery scheduling method based on graph structure perception, characterized in that: The following steps are involved: Step 1: Construct a dynamic demand map model G = (V, E), where G is a complete directed graph including nodes and edges between nodes, V represents a vertex set, including a parking lot set D, a customer set C, and a charging station set F, and E is a set of edges in the graph; at the same time, set the vehicle routing problem constraints and objective function, and build a mathematical model for the electric vehicle pickup and delivery scheduling problem; Step 2: The objective function is set as the energy consumption factor, comprehensively considering the influence of acceleration, speed, terrain, and power system efficiency on the energy consumption of electric vehicles, and embedding the above path details into the energy cost coefficient linearly related to mass to establish the energy consumption model E between any two points i and j. ij ; Step 3: Construct the vehicle routing decision Markov decision process; Step 4: Construct a neural network model for the vehicle pick-up and delivery routing problem based on the decision-making process; The policy network consists of an encoder and a decoder. We innovatively added a graph attention mechanism to the encoder, and we also added an adjacency matrix and feature embedding module to process node features. The network processes the input current state and outputs the probability of node selection; The baseline network uses a similar structure to the policy network and uses a greedy strategy to select the action with the highest probability, reducing the variance in the policy network and improving the stability and efficiency of training; Step 5: Design a Reinforcement Learning algorithm to train the network model in Step 4.

2. According to the graph structure perception-based learning electric vehicle pickup and delivery scheduling method of claim 1, it is characterized in that: In step 1, the mathematical model of the problem is: τ0=0(7) u0=0(10) where x ij As a binary decision variable, when a vehicle starts from node i and visits node j, x ij =1, otherwise 0; q j represents the demand of customer point j; d ij represents the Euclidean distance between node i and node j; v ij represents the average speed of the vehicle between node i and node j; t ij represents the travel time of the vehicle between node i and node j; E ij represents the energy consumption of the vehicle between node i and node j; u ij represents the load of the vehicle between node i and node j; U represents the maximum capacity of the vehicle; Q represents the battery capacity of the vehicle; g j represents the service time of the vehicle arriving at node j; e j represents the remaining power of the vehicle after it reaches node j; τ j represents the time required for the vehicle to reach node j; The objective function (1) of the model indicates that the driving cost of electric vehicles is minimized under the constraints of constraints (2) to (18); where constraint (2) indicates that each customer visits the vehicle only once; constraint (3) indicates that not all charging stations need to be visited; constraint (4) indicates flow conservation; constraint (5) indicates that the number of available vehicles does not exceed K; constraint (6) indicates that the vertices are sorted according to the time when the vehicles visit them; constraints (7) and (8) indicate that the electric vehicles start from the parking lot at the beginning and the driving time of a single vehicle cannot exceed the maximum time constraint; constraint (9) indicates the real-time cargo load of the electric vehicle; constraints (10) and (11) indicates that the electric vehicle is unloaded when leaving the garage and the vehicle load cannot exceed the maximum load; constraints (12) and (13) indicate tracking the vehicle SOC status and ensuring that the battery charge is within the normal range; constraint (14) indicates that the vehicle leaves the garage in a fully charged state and is fully charged when visiting the charging station; constraint (15) indicates the decision variable constraint; constraint (16) indicates the calculation of the arrival time of two adjacent nodes in the path; constraint (17) is a priority constraint, which ensures that the arrival time of the pickup point is earlier than the arrival time of the delivery node; finally, constraint (18) imposes the non-negativity of the arrival time.

3. According to the graph structure perception-based learning electric vehicle pickup and delivery scheduling method of claim 1, it is characterized in that: In step 2, the mechanical power P of the vehicle on the driving section is first determined by comprehensively considering the acceleration, speed, terrain, and power system efficiency of the vehicle on all sections. ij , calculated as follows: P ij =(m ij ·a+m ij ·g·(sin(α ij )+C r ·cos(α ij )+0.5°C d ·p·A·v ij 2 )·v ij In the formula, m ij represents the total mass of the vehicle traveling on edge (i, j), a represents the acceleration of the vehicle, A represents the front area of ​​the EV, and C d is the aerodynamic drag coefficient, C r is the rolling resistance coefficient, ρ is the air density, g is the gravity constant, α ij is the average road slope between vertex i and vertex j, v ij represents the average speed between vertex i and vertex j; use t ij represents the travel time of the vehicle along (i, j), φ d and φ r They represent the power of the motor during propulsion and regenerative braking, respectively. and They represent the battery charging and discharging efficiency, vehicle energy consumption E ij The calculation is as follows: Based on the above explanation, the energy consumption cost of all road sections in the directed graph in step 1 can be calculated.

4. The learning-based electric vehicle pickup and delivery scheduling method based on graph structure perception according to claim 1 is characterized in that: In step 3, the mathematical model of the electric vehicle pickup and delivery scheduling problem is converted into a Markov decision process; Define the vehicle state S t =(x t , v t ) indicates that at time step t the system state is represented by the vertex state and the vehicle state v t =[e t , τ t ,u t ]composition; u t , e t , τ t They represent static information, dynamic information, vertex two-dimensional coordinates, current vertex demand, load, power status, and driving time respectively; Determine the node to be visited by the vehicle in the current state, and the feasible action a t , the final generated action sequence {a0, a1, ..., a T }. Considering the dynamic requirements, the state transfer equation from node i to node j is established, and the next state S t+1 According to the current state S t Update, when executing action a t Then update the vehicle status. The reward function is expressed as where r t It is expressed as the negative value of the incremental travel energy consumption at a step t.

5. The learning-based electric vehicle pickup and delivery scheduling method based on graph structure perception according to claim 1 is characterized in that: In step 4, the composition of the neural network model is as follows: The policy network consists of an encoder and a decoder, and we add an adjacency matrix and feature embedding module to process node features. Given the graph information, the parameterized policy is expressed as: The graph attention mechanism learns graph structure information in the encoder, transfers graph information through the adjacency matrix, reduces feature calculation, and trains a weight matrix W∈R for each node after a given input. F×F′ , then LeakyReLU is used to provide nonlinearity and normalized using the softmax function, the graph attention layer calculates the attention coefficient for the node pair (i, j), and then updates its own features as input by calculating the linear combination of the expected adjacent node features. The following is the calculation equation group: In order to stabilize the learning process, a multi-head attention mechanism is added to average the output using a nonlinear function, as shown in the following formula: The adjacency matrix and feature embedding module jointly process node feature information. The adjacency matrix is ​​used to filter out valid nodes that can be reached. The feature embedding module will process the vehicle state S t And the routing feature R t Perform linear projection and maximum pooling operations, and then perform feature fusion to obtain the output feature embedding O t ,Then the feature embedding step will concatenate the vehicle state features and routing features and pass them to the decoder; The decoder generates a probability vector for selecting vertices, which is obtained by the graph embedding, the embedding at time step t-1, and the feature embedding O t The context vector composed of And the multi-head attention calculation outputs a probability vector at each time step to represent the probability of each node being selected. The specific calculation group is shown below Computational Context Glimpses in are all trainable matrices. and k t =W K h N , so the single-head attention h g The compatibility with all nodes in step t is calculated as follows: Finally, we use the softmax function to calculate the probability vector and output the selection probability of the vertex, which is expressed as follows: p t =softmax(h t +Z·Mask t ) Where Mask t Represents the mask rule, if the vertex is masked, it is 1, otherwise it is 0; The baseline network uses a similar structure to the policy network and uses a greedy strategy to select the action with the highest probability, reducing the variance in the policy network and improving the stability and efficiency of training.

6. The learning-based electric vehicle pickup and delivery scheduling method based on graph structure perception according to claim 1 is characterized in that: In step 5, the training process of the Reinforce reinforcement learning algorithm is as follows: The Reinforce reinforcement learning algorithm is used to train the model. The training goal is to find the policy parameter θ that maximizes the reward. The target training function can be expressed as: In the formula, R(π θ ) represents the total accumulated reward, p θ (a|s) represents the probability of selecting a node, Represents the log gradient of the probability of choosing an action.

7. The learning-based electric vehicle pickup and delivery scheduling method based on graph structure perception according to claim 1 is characterized in that: The examples in the training described above are all generated data. In each time step, the Actor network will generate probabilities and update the status until the end. During the learning process, the model will update the network parameters according to the gradient. The graph attention mechanism focuses on the graph structure information to reduce the feature calculation of invalid points, and the feature fusion is performed in advance through the feature embedding module to improve the computing efficiency. The Reinforce reinforcement learning algorithm is selected to retain the Actor network and combine it with the Baseline based on policy optimization, which reduces the variance in the Actor network and improves the stability and efficiency of training.