Vehicle path planning method and system based on multi-head adaptive Actor-Critic algorithm

The vehicle path planning system built using the multi-head adaptive Actor-Critic algorithm solves the problems of poor adaptability and low solution efficiency in vehicle path planning under complex environments, and realizes fast and efficient path optimization in large-scale VRP problems.

CN120927016APending Publication Date: 2025-11-11KAILI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511050328.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies lack effective methods to efficiently seek optimal solutions for vehicle routing in complex environments. In particular, when facing large-scale vehicle routing problems, existing methods suffer from poor adaptability, long solution times, and degraded solution quality.

Method used

A vehicle path planning method based on the multi-head adaptive Actor-Critic algorithm is adopted. The urban road network data is mapped into a multi-dimensional vehicle path coding sequence through an encoder. The MHAAC model is constructed by combining a deep reinforcement learning framework and the multi-head adaptive Actor-Critic algorithm to generate a path planning system. The path construction is optimized by dynamically adjusting the parameters of the multi-head Actor network and the Critic network.

Benefits of technology

Under the constraint of meeting customer requirements, it can quickly generate the global optimal path solution, which solves the problems of poor adaptability and long solution time caused by sparse environmental information features, fixed feature embedding and single decoding strategy, and improves the efficiency of path planning and the quality of solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120927016A_ABST
    Figure CN120927016A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle path planning method based on a multi-head adaptive Actor-Critic algorithm, and the method comprises the steps: obtaining urban road network data comprising at least one warehouse center and more than two demand client nodes through a simulation generation technology, obtaining a training and testing data set, carrying out the coding processing of the urban road network node training data, and carrying out the coding processing of the urban road network node training data, obtaining a multi-dimensional vehicle path coding sequence on a two-dimensional plane space [0, 1] * [0, 1]; a deep reinforcement learning framework is built based on the multi-dimensional vehicle path coding sequence, a multi-head adaptive Actor-Critic algorithm is integrated to build an MHAAC model, and a vehicle path planning system is formed; and the constructed MHAAC model is adopted to carry out path solving on urban road network data, and a global optimal path scheme is generated under the condition that all customer demand constraint conditions are met. According to the scheme, the problems of poor adaptivity, prolonged solving time, reduced solution quality and the like caused by limitations of sparse environmental information features, fixed feature embedding, static parameter generation, single decoding strategy and the like are successfully solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of intelligent transportation, specifically to a vehicle path planning method and system based on a multi-head adaptive Actor-Critic algorithm. Background Technology

[0002] The Vehicle Routing Problem (VRP) is a classic combinatorial optimization problem in applied mathematics and computer science. Its core objective is to construct the globally optimal transportation route to minimize costs while satisfying all customer needs (such as time window constraints and load limits). This problem has wide applications in real-world scenarios such as food delivery, ride-sharing scheduling, and logistics warehouse transportation. With the rapid growth in demand for smart cities and instant services, higher requirements are being placed on efficient and robust route planning methods.

[0003] Traditional solution methods face three main challenges: (1) Mathematical combinatorial optimization methods based on exact algorithms (such as branch and bound) can guarantee theoretical optimality, but their exponential computational complexity makes it difficult to cope with NP-hard problems of actual scale; (2) Heuristic methods (such as simulated annealing and genetic algorithms) rely on manually designed domain knowledge to reduce the search space, but their rule generalization ability is limited and they need to be repeatedly tuned for different scenarios; (3) Deep learning methods (such as Pointer Networks based on attention mechanisms) automatically learn node feature representations through encoder-decoder architecture and combine reinforcement learning to train policy networks, which significantly improves the solution efficiency of medium-sized problems. However, existing neural heuristic algorithms still have some limitations: ① When the problem scale increases, the static graph embedding method will lead to a decrease in the distinguishability of node features; ② The fixed parameters of the policy network in the Actor-Critic framework are difficult to adapt to dynamic environmental changes (such as real-time traffic changes); ③ The greedy search strategy in the decoding stage is prone to getting trapped in local optima.

[0004] In summary, existing methods lack an effective technical means to efficiently seek the optimal solution for vehicle routing in complex environments. Summary of the Invention

[0005] To address the problems existing in the prior art, this invention provides a vehicle path planning method and system based on the multi-head adaptive Actor-Critic algorithm, in order to solve the technical problem of lacking an efficient solution for vehicle path planning in complex environments.

[0006] This invention provides a vehicle path planning method based on a multi-head adaptive Actor-Critic algorithm, comprising:

[0007] S1. Obtain urban road network data, which includes at least one warehouse center, two customer geographic location coordinates, and demand values;

[0008] S2. The customer's geographical location coordinates are mapped to a 128-dimensional static feature vector through a 1D convolutional layer of the encoder, the demand value is encoded into a dynamic feature vector, and a multi-dimensional vehicle path encoding sequence on a two-dimensional plane is obtained.

[0009] S3. Based on the multi-dimensional vehicle path coding sequence, a deep reinforcement learning framework is built, and the multi-head adaptive Actor-Critic algorithm is integrated to construct the MHAAC model (the full name of the MHAAC model in this invention is Multi-Head Adaptive Actor-Critic), forming a vehicle path planning system.

[0010] S4. Based on all customer requirement constraints, generate feasible route planning solutions through the vehicle route planning system.

[0011] Optionally, the acquisition of urban road network data includes at least one warehouse center, two customer geographic location coordinates, and demand values, including:

[0012] S101. Randomly sample customer coordinates s in the interval [0,1] using uniform distribution. i =(x i ,y i )∈[0,1]×[0,1], and set the last coordinate as the warehouse center s0=(x0,y0)∈[0,1]×[0,1];

[0013] S102. Randomly generate a positive integer demand value d within a preset range. i ∈R + ;

[0014] S103. Training and test data generators are designed based on the generation principles of S101 and S102.

[0015] Optionally, the customer's geographic location coordinates are mapped to a 128-dimensional static feature vector through a 1D convolutional layer of the encoder, the demand value is encoded into a dynamic feature vector, and a multi-dimensional vehicle path encoding sequence in a two-dimensional planar space is obtained, including:

[0016] S201. Format the training data as a view G = (V, E), where V = {x} i |x i =(s i ,d i Let {i = 0, 1, 2, ..., n} represent the set of nodes, and E represent the set of edges between nodes;

[0017] S202. The test dataset generated by the test data generator is labeled and stored using a random seed, and the training dataset x generated by the training data generator is stored. i =(s i ,d i Using a one-dimensional convolutional layer for vector mapping, a feature vector space with a coding size of 128 dimensions is obtained, which is called a multi-dimensional vehicle path coding sequence on the two-dimensional plane space [0,1]×[0,1], as follows:

[0018] Embedding = Embedding Layer(s) i ,d i ).

[0019] Optionally, a deep reinforcement learning framework is built based on the multi-dimensional vehicle path encoding sequence, and a multi-head adaptive Actor-Critic algorithm is integrated to construct an MHAAC model, forming a vehicle path planning system, including:

[0020] S301, Obtain the output sequence h of the multi-dimensional vehicle path encoding sequence by tensor copying. s , used as context vector c t The generation of the decoder begins with the warehouse center and initializes its state.

[0021] S302. The decoder uses a custom masking scheme based on the output sequence and the LSTM hidden state h. t Constructing dynamic decision states Using the dynamic decision state as input to the multi-head Actor network, the probability distribution of unmasked nodes is obtained;

[0022] S303. Optimize path construction using a search algorithm and optimize network parameters through reinforcement learning loops.

[0023] Optionally, the multi-dimensional vehicle path encoding sequence is used to obtain its output sequence h through tensor copying. s , used as context vector c t The generation; starting from the repository center, and initializing the decoder's state, including:

[0024] S3011, Calculate the encoder output sequence h s and the current hidden state h of the decoder t The relevance score is expressed as:

[0025] F(h t ,h s )=V·tanh(W1h t +W2h s ),

[0026] Among them, h t h represents the hidden state at the current time step t of the decoder. s This represents the output sequence of the encoder, where W1 and W2 are linear transformation matrices, and V is a trainable weight vector used to map the activated hidden state to a scalar score.

[0027] S3012. Normalize the scalar score using Softmax to obtain the attention weights, expressed as follows:

[0028]

[0029] S3013. Generate context vector c using weighted summation. t It represents the current environmental context, and the formula is:

[0030]

[0031] Optionally, the decoder uses a custom masking scheme based on the output sequence and the LSTM hidden state h. t Constructing dynamic decision states Using the dynamic decision state as input to the multi-head Actor network, the probability distribution of unmasked nodes is obtained, including:

[0032] S3021, The dynamic decision state is... and customer node characteristics d t As input, multiple convolutional transformations are applied, then the extracted features are distributed to multi-head parallel processing, and finally the output results are merged and operated on with the training parameters to generate probability base values. The core mathematical expression can be shown as follows:

[0033]

[0034] Among them, a t To determine the degree of correlation between the current input node information and the information in the next decoding step t, For the embedded input, h t The hidden state values ​​of the LSTM are the historical sequence information; ";" indicates the concatenation of two tensors. t For the context vector, P(y) t+1 |Y t ,X t W represents the probability of the next node. a W c v a and v c These are trainable parameters;

[0035] S3022. The decoder sets an option to block invalid or constraint-violation nodes during the decoding process, so as to block nodes that have already been accessed.

[0036] S3023. The dynamic attention parameters v of the Critic network are generated through a fully connected layer FC(·), expressed as:

[0037]

[0038] Where FC(·) represents a fully connected layer with parameters W,b and activation value σ, T is the length of the input sequence, and X is the value of σ. t For the customer node features at time step t, the dynamic attention parameter v is used as the adjustment method for the Critic network, and after attention calculation, it is passed to the multi-head Actor network to generate dynamic attention weights, expressed as:

[0039] u = v T ·tanh(q+e+d),

[0040] Where q is the query vector of the embedded features, e is the reference vector of the embedded features, d is the vector of dynamic features, and u is the generated attention scoring matrix.

[0041] Optionally, the step of optimizing path construction using a search algorithm and iteratively optimizing network parameters through reinforcement learning includes:

[0042] The decoder is divided into a node probability distribution generation part and a path construction part. The path construction part is further divided into an initial stage, an intermediate stage, and a later stage. Each stage employs a different search strategy to select the next node using the node probability distribution and adds the selected node to the partial sequence solution until the complete solution is output. Finally, let θ and φ be the weight vectors of the Actor network and the Critic network, respectively. θ and φ are correlated, and the reinforcement learning loop optimization is represented as:

[0043]

[0044] Where the superscript indicates the variable of the nth instance, and N represents the number of problem samples. R is the approximate reward calculated for the Critic network, and R is the actual reward value. Greedy search is used in the initial stage, random search is used in the middle stage, and greedy search is used in the later stage.

[0045] This invention also provides a vehicle path planning system based on a multi-head adaptive Actor-Critic algorithm, comprising:

[0046] The data generation module is used to acquire urban road network data, which includes at least one warehouse center, two customer geographic location coordinates, and demand values.

[0047] The data processing module is used to map the customer's geographic location coordinates into a 128-dimensional static feature vector through a 1D convolutional layer of the encoder, encode the demand value into a dynamic feature vector, and obtain a multi-dimensional vehicle path encoding sequence in a two-dimensional planar space.

[0048] The path planning module is used to build a deep reinforcement learning framework based on the multi-dimensional vehicle path encoding sequence, integrate the multi-head adaptive Actor-Critic algorithm to construct the MHAAC model, and form a vehicle path planning system.

[0049] The route generation module is used to generate feasible route planning schemes through the vehicle route planning system based on all customer requirement constraints.

[0050] Compared with the prior art, the present invention:

[0051] By employing simulation generation technology, urban road network data containing at least one warehouse center and two demand customer nodes is acquired. This urban road network node data is then processed to obtain a multi-dimensional vehicle path coding sequence on a two-dimensional plane space [0,1]×[0,1], which is used for model training. Based on this multi-dimensional vehicle path coding sequence, a deep reinforcement learning framework is built, and a multi-head adaptive Actor-Critic algorithm is integrated to construct the MHAAC model, forming a vehicle path planning system. The constructed MHAAC model is used to solve paths on VRP datasets (including program-generated and publicly available VRP datasets), achieving the generation of globally optimal path solutions while satisfying all customer demand constraints. This solution successfully solves the problems of poor adaptability, prolonged solution time, and decreased solution quality caused by limitations such as sparse environmental information features, fixed feature embedding, static parameter generation, and a single decoding strategy. Attached Figure Description

[0052] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0055] Figure 2 This is a schematic diagram of data map matching provided in one embodiment of the present invention;

[0056] Figure 3 This is a schematic diagram of the path data topology structure according to an embodiment of the present invention;

[0057] Figure 4 This is a flowchart of the system framework based on the MHAAC model according to an embodiment of the present invention;

[0058] Figure 5 This is a schematic diagram of the path search and selection process according to an embodiment of the present invention. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other implementation cases obtained by those skilled in the art without creative effort are within the scope of protection of this application. Functional units with the same reference numerals in the examples of this invention have the same or similar structures and functions.

[0060] In one embodiment, reference is made to the appendix. Figure 1 As shown, a vehicle path planning method based on the multi-head adaptive Actor-Critic algorithm is provided, including the following steps:

[0061] S1. Obtain urban road network data, which includes at least one warehouse center, two customer geographic location coordinates, and demand values, specifically including:

[0062] S101. Randomly sample customer coordinates s in the interval [0,1] using uniform distribution. i =(x i ,y i )∈[0,1]×[0,1], and set the last coordinate as the warehouse center s0=(x0,y0)∈[0,1]×[0,1];

[0063] S102. Randomly generate a positive integer demand value d within a preset range. i ∈R + ;

[0064] S103. Design training data and test data generators based on the generation principles of S101 and S102 respectively;

[0065] S2. The customer's geographic location coordinates are mapped to a 128-dimensional static feature vector through a 1D convolutional layer of the encoder. The demand value is encoded into a dynamic feature vector, resulting in a multi-dimensional vehicle path encoding sequence in a two-dimensional planar space, specifically including:

[0066] S201. Format the generated data as a view G = (V, E), where V = {x} i |x i =(s i ,d i Let {i = 0, 1, 2, ..., n} represent the set of nodes, and E represent the set of edges between nodes.

[0067] S202. The test dataset generated by the test data generator is labeled and stored using a random seed, and the training dataset x generated by the training data generator is... i =(s i ,d i Using a one-dimensional convolutional layer for vector mapping, a feature vector space with a coding size of 128 dimensions is obtained, which is called a multi-dimensional vehicle path coding sequence on the two-dimensional plane space [0,1]×[0,1], as follows:

[0068] Embedding = Embedding Layer(s) i ,d i (1)

[0069] This involves constructing a globally optimal transportation route to minimize costs, while meeting customer needs (such as time window constraints and load limits). See [link to relevant documentation]. Figure 2 The diagram illustrates urban data matching for a real-world scenario, including a warehouse, ten customer nodes, and vehicles with capacity limitations. Each customer node has a corresponding demand value. Specifically, firstly, environmental feature analysis is performed based on the actual scenario, involving the location distribution, quantity relationships, and vehicle capacity of the warehouse and customer nodes. Based on this feature information, a simulation environment data generator is built to obtain urban road network data in a two-dimensional planar space of [0,1]×[0,1] for both training and testing. Then, the training dataset is embedded using an encoder to obtain a multi-dimensional vehicle path encoding sequence. Secondly, see... Figure 3 The diagram shows the graph topology of the urban road network dataset, from which the following constraints can be defined: all vehicle paths must begin and end at warehouse nodes; each customer node must satisfy the Exactly-Once Visitation Constraint; and warehouse nodes are allowed multiple visits. Based on these topological constraints, the path construction system uses a spatial encoding module to jointly encode geographic location coordinates and their associated demand information into a 128-dimensional feature vector. The decoder then obtains the node probability distribution, and the final path search strategy constructs solutions one by one using the node probability distribution.

[0070] S3. Based on the aforementioned multi-dimensional vehicle path encoding sequence, a deep reinforcement learning framework is built, and a multi-head adaptive Actor-Critic algorithm is integrated to construct an MHAAC model (the full name of the MHAAC model in this invention is Multi-Head Adaptive Actor-Critic), forming a vehicle path planning system, specifically including:

[0071] S301: Obtain the output sequence h of the multi-dimensional vehicle path encoding sequence by tensor copying. s Used as a dynamic context c t The generation starts with the repository node and initializes the decoder's state;

[0072] S302: The decoder uses a custom masking scheme (to mask visited nodes) based on the output sequence and the LSTM hidden state h. t Constructing dynamic decision states Using the dynamic decision state as input to the multi-head Actor network, the probability distribution of unmasked nodes is obtained;

[0073] S303: Optimizes path construction using a search algorithm and optimizes network parameters through reinforcement learning loops.

[0074] See attached document Figure 4 As shown, the system is first initialized, and the urban road network data is used as the input to the encoder to obtain a multi-dimensional vehicle path encoding sequence. This sequence is then copied to obtain a tensor form that supports constraint search. This output sequence is denoted as h. s And will calculate it with the decoder's current hidden state h. t The relevance score is calculated using the following formula:

[0075] F(h t ,h s )=V·tanh(W1h t +W2h s (2)

[0076] In equations (1) and (2), s i For customer location; d i To meet customer needs; h t h represents the hidden state at the current time step t of the decoder. s Let W1 and W2 represent the encoder's output sequence, where W1 and W2 are linear transformation matrices, and V is a trainable weight vector used to map the activated hidden states to a scalar score. Simultaneously, this scalar score is normalized using Softmax to obtain the attention weights, defined by the following formula:

[0077]

[0078] Next, a context vector c is generated using a weighted sum. t It represents the current environmental context, and the formula is:

[0079]

[0080] Secondly, the dynamic decision-making state of the input representation (multi-dimensional vehicle path coding sequence) is read. and customer node characteristics d t Multi-head attention Actor networks utilize these two components to obtain the probability distribution of the next node; a t The degree of relevance to the current input node information in the next decoding step t can be specified. Specifically, firstly, h t and d t The extracted features are input to the Actor network and subjected to multi-layer convolutional transformations. The extracted features are then divided into multiple parallel processing heads. Finally, the processed results are merged and used in conjunction with the training parameters to generate probability base values. Its core mathematical expression can be shown as follows:

[0081]

[0082] Where a t To determine the degree of correlation between the current input node information and the information in the next decoding step t, For the embedded input, h t The hidden state values ​​of the LSTM are the historical sequence information; ";" indicates the concatenation of two tensors. t For the context vector, P(y) t+1 |Y t ,X t W represents the probability of the next node. a W c v a and v c These are trainable parameters.

[0083] The actual generation method of the probability base value is based on the client node position s i Customer node characteristics d t The information h of the historical solution sequence (also referred to as the partial solution sequence in this paper) recorded by the LSTM hidden state. t It is generated as input.

[0084] Simultaneously, options for masking invalid or constraint-violation solutions are set during the subsequent decoding process to ensure the feasibility of the solutions generated at each step. Homogenized dynamic parameters based on environmental information are used as the adjustment method for the Critic network, and after attention calculation, they are passed to the multi-head attention Actor network to achieve adaptive environmental control. The Critic network is based on a single-head attention mechanism, combined with dynamically generated attention parameters to enhance its adaptive capability. Unlike traditional fixed-parameter generation methods, dynamically generated parameters can be adjusted according to the average information of environmental features, allowing the model to flexibly respond to different environmental changes.

[0085] The dynamic attention parameters v of the Critic network are generated through a fully connected layer FC(·), and are expressed as follows:

[0086]

[0087] Where FC(·) represents a fully connected layer with parameters W,b and activation value σ, T is the length of the input sequence, and X is the value of σ. t Let be the features of the customer node at time step t.

[0088] This fully connected layer can generate a dynamic attention weight vector v at each time step based on the input features, adapting to the current input state and environment. The formula for calculating the attention weights is:

[0089] u = v T ·tanh(q+e+d), (9)

[0090] Where q is the embedded query vector, e is the reference vector, d is the vector of dynamic features (i.e., customer needs), and u is the generated attention scoring matrix, which determines which node to select during the decoding process.

[0091] Furthermore, firstly, the initial mask blocks all cities outside the starting point; secondly, the MHAAC model calculates that nodes A and B belong to the same score, and their probabilities will be dynamically balanced; finally, when only one unvisited city remains, the mask automatically removes its block.

[0092] See Figure 5 This describes the process from encoder to decoder and then to a hybrid search strategy (this scheme employs three solution search strategies: Greedy, Beam Search, and the designed hybrid search strategy). Specifically, first, urban road network data is input into the encoder to generate a multi-dimensional vehicle path encoding sequence. Second, leveraging multiple search strategies, tensor copying is performed on this encoding sequence (this operation primarily targets the bundle width in Beam Search), thereby obtaining the output sequence h. s Subsequently, additive attention mechanisms were used to study h. sAfter processing, the context vector c is obtained. t Next, the decoder state is initialized; specifically, the context vector c is initialized. t The hidden state h of LSTM in the decoder t The decoder takes a warehouse node (initial time) and an environment interface (which is essentially a class interface for Reward and Action interaction) as inputs and performs node probability calculations through a multi-head Actor network to obtain the node probability distribution. Finally, based on this node probability distribution, Greedy, BeamSearch, and a designed hybrid search strategy are used to select the next node, constructing a solution sequence sequentially. In practical applications, the solution sequence with the smallest path distance is selected as the optimal solution. The hybrid search strategy designed in this paper divides the decoding process into an initial stage, an intermediate stage, and a later stage. The parameter values ​​for the initial and intermediate stages are set to regulate this dynamic decoding process. Specifically, the initial stage uses greedy search, the intermediate stage uses random search, and the later stage uses greedy search again to optimize the flexibility and accuracy of the decoding process. The entire decoding process involves two neural networks: an Actor network with weight vector θ and a Critic network with weight vector φ. θ and φ are correlated, and the mathematical formula for the network strategy optimization principle is expressed as follows:

[0093]

[0094] Where the superscript indicates the variable of the nth instance, and N represents the number of problem samples. The approximate reward value calculated for the Critic network.

[0095] S4. Based on all customer requirement constraints, generate feasible route planning solutions through the vehicle route planning system.

[0096] Customer demand constraints include vehicle capacity limits, customer node traffic balancing, and customer node service requirements. These constraints are determined by calculating the probability distribution of nodes and identifying the nodes that need to be served at each time step. Based on the probability distribution of nodes generated by the multi-head Actor network, path order construction employs search algorithms (such as Greedy search and Stochastic search) to select interaction strategies within the node probability distribution until a complete path is constructed.

[0097] By employing simulation generation technology, urban road network data containing at least one warehouse center and two or more demand customer nodes is acquired to obtain training and testing datasets. The training data of the urban road network nodes is then encoded to obtain a multi-dimensional vehicle path encoding sequence on a two-dimensional plane space [0,1]×[0,1]. Based on this multi-dimensional vehicle path encoding sequence, a deep reinforcement learning framework is built, and a multi-head adaptive Actor-Critic algorithm is integrated to construct an MHAAC model, forming a vehicle path planning system. The constructed MHAAC model is used to construct paths for the urban road network data, achieving the generation of a globally optimal path solution while satisfying all customer demand constraints. This solution successfully solves the problems of poor adaptability, prolonged solution time, and decreased solution quality caused by limitations such as sparse environmental information features, fixed feature embedding, static parameter generation, and a single decoding strategy, ensuring that the optimal solution for vehicle paths is obtained quickly and accurately in complex environments.

[0098] This invention also provides a vehicle path planning system based on a multi-head adaptive Actor-Critic algorithm, comprising:

[0099] The data generation module is used to acquire urban road network data that includes at least one warehouse center and two demand customer nodes;

[0100] The data processing module is used to map the customer's geographic location coordinates in the urban road network data into static feature vectors, encode the demand data into dynamic feature vectors, and construct a multi-dimensional vehicle path coding sequence through an encoder.

[0101] The path planning module is used to build a deep reinforcement learning framework based on the vehicle path encoding sequence of the aforementioned dimension, integrate the multi-head adaptive Actor-Critic algorithm to construct the MHAAC model, and form a vehicle path planning system.

[0102] The route generation module is used to generate the optimal route planning feasible solution through the vehicle route planning system based on all customer demand constraints.

[0103] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0104] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A vehicle path planning method based on a multi-head adaptive Actor-Critic algorithm, characterized in that, include: S1. Obtain urban road network data, which includes at least one warehouse center, two customer geographic location coordinates, and demand values; S2. The customer's geographical location coordinates are mapped to a 128-dimensional static feature vector through a 1D convolutional layer of the encoder, the demand value is encoded into a dynamic feature vector, and a multi-dimensional vehicle path encoding sequence on a two-dimensional plane is obtained. S3. Based on the multi-dimensional vehicle path coding sequence, a deep reinforcement learning framework is built, and a multi-head adaptive Actor-Critic algorithm is integrated to construct the MHAAC model, forming a vehicle path planning system. S4. Based on all customer demand constraints, generate feasible route planning solutions through the vehicle route planning system.

2. The vehicle routing method based on the multi-head adaptive Actor-Critic algorithm as described in claim 1, characterized in S1, wherein the acquisition of urban road network data includes at least one warehouse center, two customer geographic location coordinates, and demand values, including: S101. Randomly sample customer coordinates s in the interval [0,1] using uniformly distributed sampling. i =(x i ,y i )∈[0,1]×[0,1], and set the last coordinate as the warehouse center s0=(x0,y0)∈[0,1]×[0,1]; S102. Randomly generate a positive integer demand value d within a preset range. i ∈R + ; S103. Training and test data generators are designed based on the generation principles of S101 and S102.

3. The vehicle routing method based on the multi-head adaptive Actor-Critic algorithm as described in claim 2, characterized in S2, is that the customer's geographic location coordinates are mapped to a 128-dimensional static feature vector through a 1D convolutional layer of the encoder, the demand value is encoded into a dynamic feature vector, and a multi-dimensional vehicle routing encoding sequence in a two-dimensional planar space is obtained, including: S201. Format the training data as a view G = (V, E), where V = {x} i |x i =(s i ,d i Let {i = 0, 1, 2, ..., n} represent the set of nodes, and E represent the set of edges between nodes; S202. The test dataset generated by the test data generator is labeled and stored using a random seed, and the training dataset x generated by the training data generator is... i =(s i ,d i Using a one-dimensional convolutional layer for vector mapping, a feature vector space with a coding size of 128 dimensions is obtained, which is called a multi-dimensional vehicle path coding sequence on the two-dimensional plane space [0,1]×[0,1], as follows: Embedding=Embedding Layer(s i ,d i )。 4. The vehicle path planning method based on the multi-head adaptive Actor-Critic algorithm as described in claim 1, characterized in S3, is that... A deep reinforcement learning framework is built based on the aforementioned multi-dimensional vehicle path encoding sequence. A multi-head adaptive Actor-Critic algorithm is integrated to construct the MHAAC model, forming a vehicle path planning system, including: S301, Obtain the output sequence h of the multi-dimensional vehicle path encoding sequence by tensor copying. s , used as context vector c t The generation starts from the warehouse center and initializes the decoder's state; S302. The decoder uses a custom masking scheme based on the output sequence and the LSTM hidden state h. t Constructing dynamic decision states Using the dynamic decision state as input to the multi-head Actor network, the probability distribution of unmasked nodes is obtained; S303. Optimize path construction using a search algorithm and optimize network parameters through reinforcement learning loops.

5. The vehicle path planning method based on the multi-head adaptive Actor-Critic algorithm as described in claim 4, characterized in S301, is that... The multi-dimensional vehicle path encoding sequence is obtained by tensor copying to obtain its output sequence h. s , used as context vector c t The generation of; Starting from the warehouse center, initialize the decoder's state, including: S3011, Calculate the encoder output sequence h s and decoder current hidden state h t The relevance score is expressed as: F(h t ,h s )=V·tanh(W1h t +W2h s ), Among them, h t h represents the hidden state at the current time step t of the decoder. s This represents the output sequence of the encoder, where W1 and W2 are linear transformation matrices, and V is a trainable weight vector used to map the activated hidden state to a scalar score. S3012. Normalize the scalar score using Softmax to obtain the attention weights, expressed as follows: S3013. Generate context vector c using weighted summation. t , representing the current context, is expressed by the formula:

6. The vehicle path planning method based on the multi-head adaptive Actor-Critic algorithm as described in claim 4, characterized in S302, wherein the decoder uses a custom masking scheme based on the output sequence and the LSTM hidden state h. t Constructing dynamic decision states Using the dynamic decision state as input to the multi-head Actor network, the probability distribution of unmasked nodes is obtained, including: S3021, The dynamic decision state is... and customer node characteristics d t The extracted features are input to a multi-head Actor network and undergo multi-layer convolutional transformations. The extracted features are then distributed across multiple heads for parallel processing. Finally, the outputs are merged and combined with the training parameters to generate probability base values. The core mathematical expression can be shown as follows: Among them, a t To determine the degree of correlation between the current input node information and the information in the next decoding step t, For the embedded input, h t The hidden state values ​​of the LSTM are the historical sequence information. ";" indicates the concatenation of two tensors. c t For the context vector, P(y) t+1 |Y t ,X t W represents the probability of the next node. a W c v a and v c These are trainable parameters; S3022. The decoder sets an option to block invalid or constraint-violation nodes during the decoding process, so as to block nodes that have already been accessed. S3023. The dynamic attention parameters v of the Critic network are generated through a fully connected layer FC(·), expressed as: Where FC(·) represents a fully connected layer with parameters W,b and activation value σ, T is the length of the input sequence, and X is the value of σ. t For the customer node features at time step t, the dynamic attention parameter v is used as the adjustment method for the Critic network, and after attention calculation, it is passed to the multi-head Actor network to generate dynamic attention weights, expressed as: u=v T ·tanh(q+e+d), Where q is the query vector of the embedded features, e is the reference vector of the embedded features, d is the vector of dynamic features, and u is the generated attention score matrix.

7. The vehicle path planning method based on the multi-head adaptive Actor-Critic algorithm as described in claim 4, characterized in that, step S303, optimizing the path construction using a search algorithm and iteratively optimizing the network parameters through reinforcement learning, includes: The decoder is divided into a node probability distribution generation part and a path construction part. The path construction part is further divided into an initial stage, an intermediate stage, and a later stage. Each stage employs a different search strategy to select the next node using the node probability distribution and add the selected node to the partial sequence solution until the complete solution is output. Finally, let θ and φ be the weight vectors of the Actor network and the Critic network, respectively. θ and φ are correlated, and the reinforcement learning loop optimization is expressed as follows: Where the superscript indicates the variable of the nth instance, and N represents the number of problem samples. R is the approximate reward calculated for the Critic network, and R is the actual reward value. Greedy search is used in the initial stage, random search is used in the middle stage, and greedy search is used in the later stage.

8. A vehicle path planning system based on a multi-head adaptive Actor-Critic algorithm, characterized in that, include: The data generation module is used to acquire urban road network data, which includes at least one warehouse center, two customer geographic location coordinates, and demand values. The data processing module is used to map the customer's geographic location coordinates into a 128-dimensional static feature vector through a 1D convolutional layer of the encoder, encode the demand value into a dynamic feature vector, and obtain a multi-dimensional vehicle path encoding sequence in a two-dimensional planar space. The path planning module is used to build a deep reinforcement learning framework based on the multi-dimensional vehicle path encoding sequence, integrate the multi-head adaptive Actor-Critic algorithm to construct the MHAAC model, and form a vehicle path planning system. The route generation module is used to generate feasible route planning schemes through the vehicle route planning system based on all customer requirement constraints.

Citation Information

Cited By

  • Method and device for measuring and calculating predicted driving duration of wharf vehicle

    CN121543856A

  • A method and device for measuring and calculating the expected driving time of a port vehicle

    CN121543856B