Route optimization method, system and equipment for large-scale electric vehicle fleet operation

By adopting a path optimization method based on graph structure and reinforcement learning in large-scale electric vehicle fleet operations, the problem that the existing technology is difficult to effectively solve the routing problem of large-scale electric vehicle with time window is achieved, and efficient and high-quality path planning is achieved.

CN119990489AActive Publication Date: 2025-05-13SOUTHEAST UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510053599.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

Existing heuristics and precise branch-price-cutting algorithms cannot be effectively solved when dealing with large-scale electric vehicle routing problems with time windows, resulting in a decrease in solution quality and efficiency.

Method used

Using a path optimization method based on graph structure and reinforcement learning, the attention model is trained to find the shortest path of the electric vehicle fleet by constructing an attention model and combining the graph embedding component, combining the random sampling method and the rollout baseline strategy gradient algorithm.

Benefits of technology

Find the shortest path for EV fleets quickly and efficiently under customer location and time conditions, support ultra-large-scale EV fleet operation examples, taking into account quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990489A_ABST
    Figure CN119990489A_ABST
Patent Text Reader

Abstract

The invention discloses a path optimization method, system and equipment for large-scale electric vehicle fleet operation, and relates to the technical field of vehicle path planning. The method comprises the steps of defining an electric vehicle routing problem with a time window based on a graph structure, and describing the electric vehicle routing problem with the time window based on a reinforcement learning angle; constructing an attention model for aggregating local and global information of the graph structure; sampling a graph structure solving result by adopting a random sampling method; and finally, training the attention model by adopting a rollout baseline strategy gradient algorithm. According to the method, the shortest path can be quickly and efficiently searched for the electric automobile fleet and the routing decision can be made under the condition of ensuring the position and time of the customer, so that the total driving distance of the fleet is minimum, and the method has great expandability and can support the operation instance of the super-large-scale electric automobile fleet.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automobile path planning, and in particular to a path optimization method, system and equipment for large-scale electric vehicle fleet operations. Background Art

[0002] In the context of the electric vehicle routing problem with time windows, a capable fleet of electric vehicles is responsible for providing services to customers in a specific area. Each customer has his or her own needs, and these needs must be met within a specified time. Therefore, more careful route planning and more accurate time prediction are required to ensure that customer needs are met within the specified time and to find the path with the minimum total distance.

[0003] The existing heuristic algorithms (VNS / TS) and the exact branch-price-cut algorithm proposed based on four variants of the electric vehicle routing problem with time windows can provide high-quality solutions to the benchmark instance of the electric vehicle routing problem with time windows. However, the quality and efficiency of the solutions decrease as the instance size increases, and they cannot solve large-scale electric vehicle routing problem instances with time windows. Therefore, we propose a path optimization method, system and device for large-scale electric vehicle fleet operations. Summary of the invention

[0004] The purpose of the present invention is to provide a path optimization method, system and equipment for large-scale electric vehicle fleet operations, which can quickly and efficiently find the shortest path for the electric vehicle fleet and make routing decisions while ensuring the customer's location and time, so that the total distance traveled by the fleet is minimized. It can support ultra-large-scale electric vehicle fleet operation instances and take into account both quality and efficiency.

[0005] According to a first aspect of the present invention, in order to achieve the above-mentioned purpose, the present invention provides the following technical solution: a path optimization method for large-scale electric vehicle fleet operation, comprising the following steps:

[0006] The electric vehicle routing problem with time windows is defined based on the graph structure, and the system state, reward scheme and shielding strategy of the electric vehicle routing problem with time windows are described from the perspective of reinforcement learning.

[0007] Build an attention model and combine it with the graph embedding component to aggregate local and global information of the graph structure;

[0008] The random sampling method is used to sample the graph structure solution results;

[0009] The rollout baseline policy gradient algorithm is used to train the attention model, and the trained attention model is used to find the shortest path for the electric vehicle fleet.

[0010] Furthermore, the electric vehicle routing problem with time windows is defined based on the graph structure as follows:

[0011] (21) The weight of each edge of the graph is defined as the Euclidean distance between the connected vertices. The graph structure includes three types of vertices: Client V c 、Station V s and Warehouse V d ;

[0012] At this time, the solution to the electric vehicle routing problem with a time window is the vertex in the graph structure, which is the planned route of the electric vehicle;

[0013] (22) Each vertex i is associated with an array Associated, all vertex arrays form an information set X t , used to decode the local information at the vertex at step t, where x i and z i represents the geographic coordinates of vertex i, e i and l i represents the corresponding time window, represents the remaining demand of vertex i at decoding step t;

[0014] (23) Each vertex shares a set of global variables:

[0015] G t ={τ t ,b t ,σ t}

[0016] Where τ t represents the time of the active EV at the beginning of decoding step t, and the initial value is set to 0; b t Represents the battery power, and the initial value is set to the battery capacity of the electric vehicle; σ t Represents the number of available electric vehicles, and its initial value is set to the size of the fleet.

[0017] Furthermore, the system state, reward scheme and shielding strategy of the electric vehicle routing problem with time windows are described from the perspective of reinforcement learning, as follows:

[0018] (31) In the electric vehicle routing problem with time windows, the system state is the information set X in the graph structure. t and the global variable G t Representation of;

[0019] Assume that there is an agent, which will add a vertex to the end of the current sequence at each step according to the current system state and the given information, and the system state will change accordingly, repeating this process until the termination condition is met;

[0020] Assume that at step t m Terminate at the point where all customer requirements are met, and define y t represents the vertex selected at step t, and defines Y t represents the vertex sequence formed before step t;

[0021] At each decoding step t, given the set X t , global variable G t and travel history t , the probability of each vertex i being added to the sequence is estimated through the probability distribution formula, and the next vertex y to be visited is decoded according to this probability distribution t+1 , the probability distribution formula is as follows:

[0022] P(y t+1 =i|X t ,G t ,Y t )

[0023] In the formula, X t is the information set; G t is a set of global variables; Y t represents the vertex sequence formed before step t;

[0024] (32) Based on the obtained y t+1 , use the transfer function formula (1)-(4) to update the system state, as follows:

[0025] (32.1) System time τ t+1 :

[0026]

[0027] Where V c Represents customer, V s Indicates station, V d Indicates warehouse; Indicates taking the maximum value of the system time and the vertex; y t represents the vertex selected in step t, w(y t ,y t+1 ) is from vertex y t To vertex y t+1 travel time; b t Indicates the battery power, re(b t ) is the time required to fully charge the battery from a given standard, and s is a constant representing the service time of each customer vertex;

[0028] (32.2)Battery capacity b t+1 :

[0029]

[0030] Where b t Indicates battery power, V c represents the customer, y t represents the vertex selected at step t, f(y t ,y t+1 ) is the number of cells from the electric vehicle vertex y t To vertex y t+1 The energy consumption is, B is the battery capacity;

[0031] (32.3) The number of electric vehicles available at each vertex is σ t and the remaining demand ε t i renew:

[0032]

[0033] Where V d Represents the warehouse, y t represents the vertex selected at step t.

[0034] (32.4) Set the reward function, which is used to guide the attention model to generate solutions subject to conditional constraints. The first term in the reward function is set to the negative total distance traveled by the fleet to support short-distance solutions. The other terms are penalty terms for the problem constraints, as follows:

[0035]

[0036] In the formula, σ t represents the number of electric vehicles, b t Indicates the battery level, y t represents the vertex selected at step t, Y t represents the vertex sequence formed before step t, w(y t-1 ,y t ) is along the edge (y t-1 ,y t ) travel time, S Along the trajectory The number of site visits, β1, β2 and β3 are three negative constants.

[0037] Furthermore, an attention model is constructed and combined with the graph embedding component to aggregate local and global information of the graph structure, as follows:

[0038] The attention model consists of a graph embedding component, an attention mechanism, and an LSTM decoder;

[0039] (41)Graph Embedding Component

[0040] First, the input information set X of the attention model t and the global variable G t Mapped to high-dimensional vector space, respectively and And use the Structure2Vec tool to synthesize the embedding vector;

[0041] For each vertex vector i, the vector Initialize and then update using the recursive formula After p rounds of recursion, the attention network model will generate a ξ-dimensional vector for each vertex i And set for

[0042] In each round of recursion, global information and position information are aggregated through the first two terms of the recursive formula, and information from different vertices and edges is propagated to each other through the last two summation terms, so that the final embedding vector It contains both local information and global information. The recursive formula is as follows:

[0043]

[0044] Where N(i) is the set of vertices connected to vertex i by an edge, and this set is called the neighborhood of vertex i. w(i,j) represents the travel time on edge (i,j). θ1, θ2, θ3, θ4, and θ5 are trainable variables. Relu is a nonlinear activation function, and relu(x) = max{0,x}.

[0045] (42) Attention Mechanism

[0046] Based on embedding vector Utilize context-based attention mechanism to calculate the visit probability of each vertex i;

[0047] First, calculate the context vector c t , the state of the entire graph structure is specified as the weighted sum of all embedded vectors, that is:

[0048]

[0049] The weight of each vertex is defined as follows:

[0050] a t =softmax(v t ) (8)

[0051]

[0052] In the formula, a t context vector, is the vector v t The i-th item, h t is the hidden memory state of the LSTM decoder, θ v and θ u is a trainable variable, [;] means connecting the two vectors on both sides of the symbol ";", and tanh is a nonlinear activation function. is the normalized exponential function applied to the vector,

[0053]

[0054] Then, in The probability of visiting each vertex i is estimated in the following formula:

[0055] p t =softmax(p t ) (10)

[0056]

[0057] In the formula, p t context vector, is the vector g t The i-th entry of c and θ g is a trainable variable.

[0058] (43) Shielding strategy

[0059] Design the following blocking strategies to exclude infeasible routes, as follows:

[0060] Vertex j represents a customer whose unsatisfied demand is either zero or exceeds the remaining goods of the EV;

[0061] Vertex j represents a customer, and the current battery charge of the electric vehicle is b t It is impossible to support electric vehicles to complete the journey from vertex i to vertex j and then to the car factory;

[0062] The earliest time to reach vertex j violates the time window constraint, that is, τ t +w(i,j)>l j ;

[0063] The electric car travels from vertex i to vertex j and cannot return to the warehouse before the specified time T ends;

[0064] If the electric vehicle is currently in the warehouse and has no remaining goods on any customer vertex, all vertices except the warehouse are blocked;

[0065] (44) LSTM Decoder: Use LSTM to model the decoder network.

[0066] Furthermore, a random sampling method is used to sample the graph structure solution results, as follows:

[0067] For all vertices i, at each decoding step t, according to The probability distribution described samples the next vertex to be visited, repeats this process to obtain multiple solutions for an instance, and reports the solution with the shortest distance.

[0068] Furthermore, the rollout baseline policy gradient algorithm is used to train the attention model to obtain the trained attention model, which is used to find the shortest path for the electric vehicle fleet, as follows:

[0069] (61) Define the loss function

[0070] The loss function represents the use of a random strategy π θ The negative expected reward of the sampled trajectory Y is given by:

[0071]

[0072] (62) Gradient estimation

[0073] Estimate the gradient of the loss function L(θ) with respect to the trainable variable θ as follows:

[0074]

[0075] In the formula, X t is the information set; G t is a set of global variables; Y t represents the vertex sequence formed before step t; parameter N is the batch size, X [i] is the i-th training example in the batch, Y [i] To use π θ The corresponding solution generated; BL() represents the rollout baseline, P θ (Y [i] │X [i] ) represents a given training example X [i] Use a random strategy π θ Generate solution Y [i] probability.

[0076] According to a second aspect of the present invention, the present invention provides a path optimization system for large-scale electric vehicle fleet operations, which is used to implement the path optimization method for large-scale electric vehicle fleet operations, including:

[0077] A definition module is used to define the electric vehicle routing problem with time windows based on a graph structure, and to describe the system state, reward scheme, and shielding strategy of the electric vehicle routing problem with time windows from the perspective of reinforcement learning;

[0078] A building module for building an attention model, combining the attention model with the graph embedding component to aggregate local and global information of the graph structure;

[0079] The sampling module is used to sample the graph structure solution results using random sampling method;

[0080] The training output module is used to train the attention model using the rollout baseline policy gradient algorithm to obtain the trained attention model, which is used to find the shortest path for the electric vehicle fleet.

[0081] According to a third aspect of the present invention, the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, the above-mentioned path optimization method for large-scale electric vehicle fleet operations is adopted.

[0082] The present invention has at least the following beneficial effects:

[0083] 1. The present invention can quickly and efficiently find the shortest path for an electric vehicle fleet and make routing decisions while ensuring the customer's location and time, so that the total distance traveled by the fleet is minimized. It also has great scalability and can support ultra-large-scale electric vehicle fleet operation instances, taking into account both quality and efficiency.

[0084] 2. The attention model proposed in the present invention can quickly capture important information in the embedded graph when solving the electric vehicle routing problem with a time window, and the solution efficiency is higher.

[0085] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] Figure 1 A schematic diagram of a flow chart of an optimization method according to an embodiment of the present invention;

[0087] Figure 2 This is a schematic diagram of the structure of an electric vehicle routing problem with a time window in an embodiment of the present invention;

[0088] Figure 3 4 is a structural diagram of the LSTM decoder in an embodiment of the present invention. DETAILED DESCRIPTION

[0089] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0090] See also Figure 1 , Figure 2 and Figure 3 The present invention provides a technical solution: a path optimization method for large-scale electric vehicle fleet operation, comprising the following steps:

[0091] S1. Based on the graph structure, the electric vehicle routing problem with time windows is defined, and the system state, reward scheme and shielding strategy of the electric vehicle routing problem with time windows are described from the perspective of reinforcement learning, as follows:

[0092] S11. The electric vehicle routing problem with time windows is defined on a graph structure, where the weight of each edge of the graph structure is the Euclidean distance between the connected vertices;

[0093] The graph structure includes three types of vertices: Clients (V c )、Station(V s ) and warehouse (V d );

[0094] At this time, the solution to the electric vehicle routing problem with a time window is the vertex in the graph structure, which is the planned route of the electric vehicle;

[0095] Each vertex i is associated with an array Associated, all vertex arrays form an information set X t , used to decode the local information at the vertex at step t, where x i and z i represents the geographic coordinates of vertex i, e i and l i represents the corresponding time window, represents the remaining demand of vertex i at decoding step t; each vertex shares a set of global variables, all of which may change over time:

[0096] G t ={τ t ,b t ,σ t}

[0097] Where τ t represents the time of the active EV (electric vehicle) at the beginning of decoding step t, and the initial value is set to 0; b tRepresents the battery power, and the initial value is set to the battery capacity of the electric vehicle; σ t represents the number of available electric vehicles, and its initial value is set to the size of the fleet;

[0098] S12. Based on the reinforcement learning perspective, the system state, reward scheme and shielding strategy of the electric vehicle routing problem with time windows are described as follows:

[0099] In the electric vehicle routing problem with time windows, the system state is the information set X in the graph structure. t and the global variable G t Representation of;

[0100] Assume that there is an agent, which will add a vertex to the end of the current sequence at each step according to the current system state and the given information, and the system state will change accordingly, repeating this process until the termination condition is met;

[0101] Assume that at step t m Terminate at the point where all customer requirements are met, and define y t represents the vertex selected at step t, and defines Y t represents the vertex sequence formed before step t;

[0102] At each decoding step t, given the set X t , global variable G t and travel history t , the probability of each vertex i being added to the sequence is estimated through the probability distribution formula, and the next vertex y to be visited is decoded according to this probability distribution t+1 , the probability distribution formula is as follows:

[0103] P(y t+1 =i|X t ,G t ,Y t )

[0104] In the formula, X t is the information set; G t is a set of global variables; Y t represents the vertex sequence formed before step t.

[0105] Based on the income t+1 , use the transfer function formula (1)-(4) to update the system state, as follows:

[0106] System time (τ t+1 ):

[0107]

[0108] Where Vc Represents customer, V s Indicates station, V d Indicates warehouse; Indicates taking the maximum value of the system time and the vertex; y t represents the vertex selected in step t, w(y t ,y t+1 ) is from vertex y t To vertex y t+1 travel time; b t Indicates the battery power, re(b t ) is the time required to fully charge the battery from a given standard, and s is a constant representing the service time of each customer vertex;

[0109] Battery power (b t+1 ):

[0110]

[0111] Where b t Indicates battery power, V c represents the customer, y t represents the vertex selected at step t, f(y t ,y t+1 ) is the number of cells from the electric vehicle vertex y t To vertex y t+1 The energy consumption is, B is the battery capacity;

[0112] The number of electric vehicles that can be used at each vertex is σ t and remaining demand renew:

[0113]

[0114] In the formula, V d Represents the warehouse, y t represents the vertex selected at step t;

[0115] Set the reward function, which is used to guide the attention model to generate solutions subject to conditional constraints. The first term in the reward function is set to the negative total distance traveled by the fleet to support short-distance solutions. The other terms are penalty terms for the problem constraints, as follows:

[0116]

[0117] In the formula, σ t represents the number of electric vehicles, b t Indicates the battery level, y t represents the vertex selected at step t, Y t represents the vertex sequence formed before step t, w(yt-1 ,y t ) is along the edge (y t-1 ,y t ) travel time, S Along the trajectory The number of site visits, β1, β2 and β3 are three negative constants.

[0118] S2. Build an attention model and combine the attention model with the graph embedding component to aggregate local and global information of the graph structure, as follows:

[0119] The attention model consists of a graph embedding component, an attention mechanism, and an LSTM decoder;

[0120] S21. Graph Embedding Component

[0121] First, the input information set X of the attention model t and the global variable G t Mapped to high-dimensional vector space, respectively and And use the Structure2Vec tool to synthesize the embedding vector;

[0122] For each vertex vector i, the vector Initialize and then update using the recursive formula After p rounds of recursion, the attention network model will generate a ξ-dimensional vector for each vertex i And set for

[0123] In each round of recursion, global information and position information are aggregated through the first two terms of the recursive formula, and information from different vertices and edges is propagated to each other through the last two summation terms, so that the final embedding vector It contains both local information and global information. The recursive formula is as follows:

[0124]

[0125] Where N(i) is the set of vertices connected to vertex i by an edge, and this set is called the neighborhood of vertex i. w(i,j) represents the travel time on edge (i,j). θ1, θ2, θ3, θ4, and θ5 are trainable variables. Relu is a nonlinear activation function, and relu(x) = max{0,x}.

[0126] S22. Attention Mechanism

[0127] Based on embedding vector Utilize context-based attention mechanism to calculate the visit probability of each vertex i;

[0128] First, calculate the context vector c t , the state of the entire graph structure is specified as the weighted sum of all embedded vectors, that is:

[0129]

[0130] The weight of each vertex is defined as follows:

[0131] a t =softmax(v t ) (8)

[0132]

[0133] In the formula, a t context vector, is the vector v t The i-th item, h t is the hidden memory state of the LSTM decoder, θ v and θ u is a trainable variable, [;] means connecting the two vectors on both sides of the symbol ";", and tanh is a nonlinear activation function. is the normalized exponential function applied to the vector,

[0134] Then, in The probability of visiting each vertex i is estimated in the following formula:

[0135] p t =softmax(p t ) (10)

[0136]

[0137] In the formula, p t context vector, is the vector g t The i-th entry of c and θ g is a trainable variable.

[0138] S23. Shielding strategy

[0139] In order to speed up the training process and ensure the feasibility of the solution, the following shielding strategy is designed to exclude infeasible routes, as follows:

[0140] Vertex j represents a customer whose unsatisfied demand is either zero or exceeds the remaining goods of the EV;

[0141] Vertex j represents a customer, and the current battery charge of the electric vehicle is b tIt is impossible to support electric vehicles to complete the journey from vertex i to vertex j and then to the car factory;

[0142] The earliest time to reach vertex j violates the time window constraint, that is, τ t +w(i,j)>l j ;

[0143] The electric car travels from vertex i to vertex j (if vertex j is a station, it will be charged at vertex j) and cannot return to the warehouse before the specified time T ends;

[0144] If the electric vehicle is currently in the warehouse and has no remaining goods on any customer vertex, all vertices except the warehouse are blocked;

[0145] S24.LSTM decoder: Use LSTM to model the decoder network, such as Figure 3 As shown, during decoding, LSTM receives the vector representation of the current position A of the EV and the memory state of the previous decoding step, and outputs a hidden state B, maintaining its trajectory information, namely C;

[0146] S3. Use random sampling to sample the graph structure solution results to quickly and comprehensively explore the solution space;

[0147] The solution space is thoroughly explored using random sampling. For all vertices i, at each decoding step t, the solution space is chosen according to p i t The probability distribution described is used to sample the next vertex to be visited, and this process is repeated to obtain multiple solutions for an instance, and the solution with the shortest distance is reported;

[0148] S4. The attention model is trained using the rollout baseline policy gradient algorithm to obtain the trained attention model, which is used to find the shortest path for the electric vehicle fleet, as follows:

[0149] S41. Define the loss function

[0150] The loss function represents the use of a random strategy π θ The negative expected reward of the sampled trajectory Y is given by:

[0151]

[0152] S42. Gradient estimation

[0153] Estimate the gradient of the loss function L(θ) with respect to the trainable variable θ as follows:

[0154]

[0155] In the formula, X t is the information set; Gt is a set of global variables; Y t represents the vertex sequence formed before step t; parameter N is the batch size, X [i] is the i-th training example in the batch, Y [i] To use π θ The corresponding solution generated; BL() represents the rollout baseline, P θ (Y [i] │X [i] ) represents a given training example X [i] Use a random strategy π θ Generate solution Y [i] probability;

[0156] S43. Example generation

[0157] In each training step, N training instances of the electric vehicle routing problem with time windows are randomly generated to quickly and efficiently find the shortest path for the electric vehicle fleet so that the total distance traveled by the fleet is minimized.

[0158] In summary, the present invention can capture the structure embedded in a given graph, find the best path for urban cargo delivery for electric vehicle fleets while ensuring customer demand and time information, minimize the total distance traveled by the fleet, and support ultra-large-scale electric vehicle fleet operation instances, taking into account both quality and efficiency.

[0159] Embodiment 2:

[0160] This embodiment provides a path optimization system for large-scale electric vehicle fleet operations, which is used to implement the path optimization method for large-scale electric vehicle fleet operations described in Embodiment 1, including:

[0161] A definition module is used to define the electric vehicle routing problem with time windows based on a graph structure, and to describe the system state, reward scheme, and shielding strategy of the electric vehicle routing problem with time windows from the perspective of reinforcement learning;

[0162] A building module for building an attention model, combining the attention model with the graph embedding component to aggregate local and global information of the graph structure;

[0163] The sampling module is used to sample the graph structure solution results using random sampling method;

[0164] The training output module is used to train the attention model using the rollout baseline policy gradient algorithm to obtain the trained attention model, which is used to find the shortest path for the electric vehicle fleet.

[0165] Specifically, the above-mentioned definition module, construction module, sampling module and training output module can be embedded in a computer processing system. The computer calls the above-mentioned modules to complete the task of path optimization of the electric vehicle fleet based on the path optimization method for large-scale electric vehicle fleet operation provided above; the above-mentioned definition module, construction module, sampling module and training output module can perform operations according to the specific steps given in the path optimization method for large-scale electric vehicle fleet operation.

[0166] It should be noted that it should be understood that the division of the various modules of the above system is only the division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated, and these modules can all be implemented in the form of software calling through processing elements; they can also be all implemented in the form of hardware; some modules can also be implemented in the form of software calling through processing elements, and some modules can be implemented in the form of hardware. For example, the definition module can be a separately established processing element, or it can be integrated in a chip of the above device for implementation. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a processing element of the above device. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the hardware integrated logic circuit in the processor element or the instructions in the form of software.

[0167] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more digital singnal processors (DSP), or one or more field programmable gate arrays (FPGA). For another example, when a module is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0168] Embodiment three:

[0169] The present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, the above-mentioned path optimization method for large-scale electric vehicle fleet operations is adopted.

[0170] It should be noted that the terminal device can be a computer device such as a desktop computer, a laptop computer or a cloud server, and the terminal device includes but is not limited to a processor and a memory. For example, the terminal device can also include input and output devices, network access devices and buses.

[0171] Furthermore, the processor may adopt a central processing unit (CPU). Of course, according to actual usage, other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. may also be adopted. The general-purpose processor may adopt a microprocessor or any conventional processor, etc., and the present application does not impose any restrictions on this.

[0172] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0173] For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances. When an element is referred to as being "assembled on", "installed on", "fixed on" or "set on" another element, it can be directly on the other element or there can also be a centered element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be a centered element at the same time. The terms "vertical", "horizontal", "up", "down", "left", "right" and similar expressions used herein are only for illustrative purposes and are not intended to be the only implementation method.

[0174] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

[0175] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

Claims

1. A path optimization method for large-scale electric vehicle fleet operation, characterized in that: The following steps are involved: The electric vehicle routing problem with time windows is defined based on the graph structure, and the system state, reward scheme and shielding strategy of the electric vehicle routing problem with time windows are described from the perspective of reinforcement learning. Build an attention model and combine it with the graph embedding component to aggregate local and global information of the graph structure; The random sampling method is used to sample the graph structure solution results; The rollout baseline policy gradient algorithm is used to train the attention model, and the trained attention model is used to find the shortest path for the electric vehicle fleet.

2. The path optimization method for large-scale electric vehicle fleet operation according to claim 1, characterized in that: The electric vehicle routing problem with time windows is defined based on the graph structure as follows: (21) The weight of each edge of the graph is defined as the Euclidean distance between the connected vertices. The graph structure includes three types of vertices: Client V c 、Station V s and Warehouse V d ; At this time, the solution to the electric vehicle routing problem with a time window is the vertex in the graph structure, which is the planned route of the electric vehicle; (22) Each vertex i is associated with an array Associated, all vertex arrays form an information set X t , used to decode the local information at the vertex at step t, where x i and z i represents the geographic coordinates of vertex i, e i and l i represents the corresponding time window, represents the remaining demand of vertex i at decoding step t; (23) Each vertex shares a set of global variables: G t ={τ t ,b t ,s t } Where τ t represents the time of the active EV at the beginning of decoding step t, and the initial value is set to 0; b t Represents the battery power, and the initial value is set to the battery capacity of the electric vehicle; σ t Represents the number of available electric vehicles, and its initial value is set to the size of the fleet.

3. The path optimization method for large-scale electric vehicle fleet operation according to claim 2, characterized in that: Based on the reinforcement learning perspective, the system state, reward scheme and shielding strategy of the electric vehicle routing problem with time windows are described as follows: (31) In the electric vehicle routing problem with time windows, the system state is the information set X in the graph structure. t and the global variable G t ' Assume that there is an agent, which will add a vertex to the end of the current sequence at each step according to the current system state and the given information, and the system state will change accordingly. This process is repeated until the termination condition is met. Assume that at step t m Terminate at the point where all customer requirements are met, and define y t represents the vertex selected at step t, and defines Y t represents the vertex sequence formed before step t; At each decoding step t, given the set X t , global variable G t and travel history t , the probability of each vertex i being added to the sequence is estimated through the probability distribution formula, and the next vertex y to be visited is decoded according to this probability distribution t+1 , the probability distribution formula is as follows: P(y t+1 =i|X t ,G t ,Y t ) Where, X t is the information set; G t is a set of global variables; Y t represents the vertex sequence formed before step t; (32) Based on the obtained y t+1 , use the transfer function formula (1)-(4) to update the system state, as follows: (32.1) System time τ t+1 : Where V c Represents customer, V s Indicates station, V d represents the warehouse; max(τ t ,e yt ) means taking the maximum value of the system time and the vertex; y t represents the vertex selected in step t, w(y t ,y t+1 ) is from vertex y t To vertex y t+1 travel time; b t Indicates the battery power, re(b t ) is the time required to fully charge the battery from a given standard, and s is a constant representing the service time of each customer vertex; (32.2)Battery capacity b t+1 : Where b t Indicates battery power, V c represents the customer, y t represents the vertex selected at step t, f(y t ,y t+1 ) is the number of cells from the electric vehicle vertex y t To vertex y t+1 The energy consumption is, B is the battery capacity; (32.3) The number of electric vehicles available at each vertex is σ t and remaining demand renew: Where V d Represents the warehouse, y t represents the vertex selected at step t; (32.4) Set the reward function, which is used to guide the attention model to generate solutions subject to conditional constraints. The first term in the reward function is set to the negative total distance traveled by the fleet to support short-distance solutions. The other terms are penalty terms for the problem constraints, as follows: In the formula, σ t represents the number of electric vehicles, b t Indicates the battery level, y t represents the vertex selected at step t, Y t represents the vertex sequence formed before step t, w(y t-1 ,y t ) is along the edge (y t-1 ,y t ) travel time, S Along the trajectory The number of site visits, β1, β2 and β3 are three negative constants.

4. The path optimization method for large-scale electric vehicle fleet operation according to claim 1, characterized in that: Construct an attention model and combine it with the graph embedding component to aggregate local and global information of the graph structure, as follows: The attention model consists of a graph embedding component, an attention mechanism, and an LSTM decoder; (41)Graph Embedding Component First, the input information set X of the attention model t and the global variable G t Mapped into a high-dimensional vector space, denoted as X t and G t , and use the Structure2Vec tool to synthesize the embedding vector; For each vertex vector i, the vector Initialize and then update using the recursive formula After p rounds of recursion, the attention network model will generate a ξ-dimensional vector for each vertex i And set for In each round of recursion, global information and position information are aggregated through the first two terms of the recursive formula, and information from different vertices and edges is propagated to each other through the last two summation terms, so that the final embedding vector δ t i It contains both local information and global information. The recursive formula is as follows: Where N(i) is the set of vertices connected to vertex i by an edge, and this set is called the neighborhood of vertex i. w(i,j) represents the travel time on edge (i,j). θ1, θ2, θ3, θ4, and θ5 are trainable variables. Relu is a nonlinear activation function, and relu(x) = max{0,x}. (42) Attention Mechanism Based on embedding vector Utilize context-based attention mechanism to calculate the visit probability of each vertex i; First, calculate the context vector c t , the state of the entire graph structure is specified as the weighted sum of all embedded vectors, that is: The weight of each vertex is defined as follows: a t =softmax(v t ) (8) In the formula, a t context vector, is the vector v t The i-th item, h t is the hidden memory state of the LSTM decoder, θ v and θ u is a trainable variable, [;] means connecting the two vectors on both sides of the symbol ";", and tanh is a nonlinear activation function. is the normalized exponential function applied to the vector, Then, in The probability of visiting each vertex i is estimated in the following formula: p t =softmax(p t ) (10) In the formula, p t context vector, is the vector g t The i-th entry of c and θ g is a trainable variable. (43) Shielding strategy Design the following blocking strategies to exclude infeasible routes, as follows: Vertex j represents a customer whose unsatisfied demand is either zero or exceeds the remaining goods of the EV; Vertex j represents a customer, and the current battery charge of the electric vehicle is b t It is impossible to support electric vehicles to complete the journey from vertex i to vertex j and then to the car factory; The earliest time to reach vertex j violates the time window constraint, that is, τ t +w(i,j)>l j ; The electric car travels from vertex i to vertex j and cannot return to the warehouse before the planned time T ends; If the electric vehicle is currently in the warehouse and has no remaining goods on any customer vertex, all vertices except the warehouse are blocked; (44) LSTM Decoder: Use LSTM to model the decoder network.

5. The path optimization method for large-scale electric vehicle fleet operation according to claim 1, characterized in that: The random sampling method is used to sample the graph structure solution results, as follows: For all vertices i, at each decoding step t, according to p t i The probability distribution described samples the next vertex to be visited, repeats this process to obtain multiple solutions for an instance, and reports the solution with the shortest distance.

6. The path optimization method for large-scale electric vehicle fleet operation according to claim 1, characterized in that: The rollout baseline policy gradient algorithm is used to train the attention model to obtain the trained attention model, which is used to find the shortest path for the electric vehicle fleet, as follows: (61) Define the loss function The loss function represents the use of a random strategy π θ The negative expected reward of the sampled trajectory Y is given by: L(θ) =-E Y~πθ [r(Y)] (12) (62) Gradient estimation Estimate the gradient of the loss function L(θ) with respect to the trainable variable θ as follows: Where, X t is the information set; G t is a set of global variables; Y t represents the vertex sequence formed before step t; parameter N is the batch size, X [i] is the i-th training example in the batch, Y [i] To use π θ The corresponding solution generated; BL() represents the rollout baseline, P θ (Y [i] │X [i] ) represents a given training example X [i] Generate solution Y using random strategy πθ [i] probability.

7. A path optimization system for large-scale electric vehicle fleet operations, used to implement the path optimization method for large-scale electric vehicle fleet operations according to any one of claims 1 to 6, characterized in that: include: A definition module is used to define the electric vehicle routing problem with time windows based on a graph structure, and to describe the system state, reward scheme, and shielding strategy of the electric vehicle routing problem with time windows from the perspective of reinforcement learning; A building module for building an attention model, combining the attention model with the graph embedding component to aggregate local and global information of the graph structure; The sampling module is used to sample the graph structure solution results using random sampling method; The training output module is used to train the attention model using the rollout baseline policy gradient algorithm to obtain the trained attention model, which is used to find the shortest path for the electric vehicle fleet.

8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, the path optimization method for large-scale electric vehicle fleet operation according to any one of claims 1 to 6 is adopted.

Citation Information

Patent Citations

  • Shared electric vehicle scheduling method based on deep reinforcement learning

    CN114971251A

  • Multi-constraint vehicle path planning method based on attention mechanism and deep reinforcement learning

    CN115759915A

  • Fresh food delivery vehicle path planning method based on multi-agent deep reinforcement learning

    CN118153789A

  • Heterogeneous vehicle type vehicle path planning method and system based on deep reinforcement learning

    CN118608021A

  • Personalized interactive semantic parsing using a graph-to-sequence model

    US20200104366A1