Vehicle path planning method and system with traffic state based on double cooperative Transformers
Through the dual-coordinated Transformer model, a ladder reward mechanism and a sine function update mechanism are introduced, and combined with the GRU gating mechanism, the problem of dynamic changes in traffic state in vehicle path planning is solved, achieving more efficient path optimization and time cost reduction.
Patent Information
- Application Number
- CN202510651454.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-25
AI Technical Summary
The existing vehicle path planning methods fail to effectively consider the dynamic changes in traffic state, resulting in a decrease in optimization capabilities in complex traffic environments and the inability to effectively reduce the total travel time.
Using a dual-collaborative Transformer method, a ladder reward mechanism and a sine function traffic state update mechanism are introduced by jointly encoding traffic state and node coordinates, and combined with the GRU gating mechanism, the model's ability to capture feature information is enhanced and path planning is optimized.
The path planning accuracy and generalization ability of the model to deal with traffic state changes is improved, the total travel time is significantly reduced, and the quality and diversity of path planning is improved.
Smart Images

Figure CN120368997A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of deep reinforcement learning and combinatorial optimization, and particularly relates to a vehicle path planning method and system with traffic states based on a dual-collaborative Transformer. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] The Traveling Salesman Problem (TSP) is a classical combinatorial optimization problem, whose goal is to find the shortest path such that a salesman can visit each given city once and only once, and finally return to the starting city.
[0004] In traditional travel planning problems, attention is usually focused on minimizing the total travel time by finding the shortest path given a certain number of scenic spots. However, this method ignores multiple real-world factors that affect the actual travel time, such as traffic conditions, environmental changes, and emergencies. Especially in practical applications, the impact of traffic conditions may cause significant time delays during the travel. Without considering traffic conditions, relying solely on the shortest path planning may lead to a substantial increase in the total travel time, thereby increasing the time cost. Therefore, it is particularly important to consider path planning problems under specific constraints, especially path optimization problems incorporating traffic condition factors.
[0005] The Traveling Salesman Problem with Traffic States (TSPTS) takes into account the dynamic changes of traffic states. Based on the original traveling salesman problem, a corresponding traffic state is added to each arriving city, and the traffic state affects the arrival travel time. Its goal is to minimize the total travel time considering both spatial distance and traffic impact. Existing technical solutions only consider minimizing the total travel path, assuming that the traffic state is constant, which leads to a significant decline in its optimization ability when facing different traffic states at various locations. Summary of the Invention To overcome the deficiencies of the above-mentioned prior art, the present invention provides a vehicle path planning method and system with traffic states based on a dual-collaborative Transformer. By jointly encoding traffic states and node coordinates, it captures the complex dynamic characteristics in path selection; introduces a stepped reward mechanism to promote the intelligent agent to deeply explore the solution space and improve the quality and diversity of solutions; designs a traffic state update mechanism based on the sine function to simulate the dynamic changes of traffic states and improve the generalization ability of the model; and combines the GRU gating mechanism to enhance the model's ability to capture feature information.
[0006] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions: The first aspect of the present invention provides a vehicle path planning method with traffic status based on dual collaborative Transformers; The vehicle path planning method with traffic status based on dual collaborative Transformers includes: Construct a set of cities with traffic status, obtain the distance and time cost between cities, construct an objective function with the goal of minimizing the total time cost, and generate an initial solution for vehicle path planning; Optimize the initial solution of the vehicle path planning using the Markov decision process, input the optimized initial solution into the trained dual collaborative Transformer model, and output the optimal path under changing traffic status; Among them, the dual collaborative Transformer model represents the feature information of each node through node feature embedding, encodes the position of the node in the path through position feature embedding, and uses the cross-aspect reference attention mechanism to integrate the spatial position and traffic status features of the node, and outputs the optimal path under changing traffic status.
[0007] As a further technical solution, the process of constructing a set of cities with traffic status, obtaining the distance and time cost between cities, constructing an objective function with the goal of minimizing the total time cost, and obtaining the initial solution of vehicle path planning is as follows: Let the set of cities be , where is the number of cities; each city , has coordinates ; each city has a traffic status , and , the traffic status reflects the congestion degree of reaching city j. At the same time, let represent the speed of reaching city i; Calculate the Euclidean distance between city i and city j. Based on the Euclidean distance, calculate the time cost from city to city as: ; In the formula, is the distance from city to city ; is the traffic status of reaching city ; Introduce a binary decision variable , used to describe whether to select the path from city i to city j, The definition is as follows: ; Construct an objective function to minimize the total time cost, expressed as: .
[0008] As a further technical solution, the process of optimizing the initial solution of the vehicle path planning by using the Markov decision process is as follows: For each instance containing nodes, the state describes the current solution , which includes the position characteristics of each node. Among them, the node characteristics are jointly derived from the node coordinates and the traffic state, that is: , where ; This state not only reflects the individual information of each node but also reflects the dynamic changes of the entire system, providing comprehensive input features for subsequent action selection; Each action represents that is adjusted by paired operators, and the reward function is defined as: ; ; ; Among them, An amplitude of the positive bias, represents the indicator function, represents the improved threshold, is the optimal solution found so far at the current state up to the time step ; The policy is parameterized using the IDACT model with parameters. At each time step t, the action (i, j) is obtained by sampling the random policy during training and inference. The policy generates appropriate action decisions by learning the connection between state-action pairs; The next state is derived from the current state , and is obtained by adjusting the given node pair using paired operators.
[0009] As a further technical solution, the dual collaborative Transformer model includes an IDAC encoder and an IDAC decoder; In the IDAC encoder, the model processes information by calculating the self-attention correlation of each aspect respectively; In the IDAC decoder, the model collects action allocation suggestions from two aspects and synthesizes them to generate the final decision.
[0010] As a further technical solution, the IDAC encoder is composed of stacked IDAC coding layers, and each coding layer includes an interaction part of node feature coding and position feature coding; in each IDAC encoder, relatively independent coding streams for node feature embedding and position feature embedding are maintained.
[0011] As a further technical solution, in the IDAC decoder, the two groups of embeddings obtained after being processed by the encoder first pass through a max-pooling sublayer and then enter a multi-head compatibility sublayer to independently generate diverse node pair selection suggestions from their respective perspectives. These suggestions are then aggregated through a feed-forward aggregation sublayer to finally generate an optimal path.
[0012] As a further technical solution, a reinforcement learning method based on proximal policy optimization is also adopted and combined with a curriculum learning strategy to improve sample efficiency.
[0013] The second aspect of the present invention provides a vehicle path planning system with traffic states based on dual collaborative Transformers.
[0014] The vehicle path planning system with traffic states based on dual collaborative Transformers includes: An initial solution acquisition module, configured to: construct a set of cities with traffic states, obtain the distance and time costs between cities, construct an objective function with the goal of minimizing the total time cost, and obtain an initial solution for vehicle path planning; An initial solution optimization module, configured to: optimize the initial solution of the vehicle path planning by using a Markov decision process. An optimal path acquisition module, configured to: input the optimized initial solution into a trained dual collaborative Transformer model and output an optimal path under changing traffic states. Among them, the dual collaborative Transformer model represents the feature information of each node through node feature embedding, encodes the position of the node in the path through position feature embedding, and uses a cross-aspect reference attention mechanism to integrate the spatial position and traffic state features of the node, and outputs an optimal path under changing traffic states.
[0015] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the steps in the vehicle path planning method with traffic states based on dual collaborative Transformers as described in the first aspect of the present invention.
[0016] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the steps in the vehicle path planning method with traffic status based on the dual-cooperative Transformer as described in the first aspect of the present invention are implemented.
[0017] The above one or more technical solutions have the following beneficial effects: Through integrating the dynamic traffic status update mechanism, multi-dimensional feature joint encoding, and the improvement of critical components, the IDACT model shows significant performance enhancement in path planning considering traffic status, has higher cost-quality and solution efficiency when dealing with path environments involving traffic status, and effectively overcomes the limitations of traditional models in such scenarios. The improvement of IDACT enables it to adapt to various traffic conditions, thus achieving better global optimization and robustness in various complex environments.
[0018] Advantages of additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0020] Figure 1 It is a flowchart of the method for the first embodiment.
[0021] Figure 2 It is a schematic diagram of the process method of the MDP in the first embodiment.
[0022] Figure 3 It is a schematic diagram of the IDAC structure in the first embodiment.
[0023] Figure 4 It is a schematic diagram of the PPO method in the first embodiment.
[0024] Figure 5 It is a schematic diagram of the comparison results of TSPTS and various models in the first embodiment.
[0025] Figure 6 It is a system structure diagram of the second embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0027] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.
[0028] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0029] The present invention captures the complex dynamic characteristics in path selection by jointly encoding traffic states and node coordinates; introduces a stepped reward mechanism to promote the intelligent agent to deeply explore the solution space and improve the quality and diversity of solutions; designs a traffic state update mechanism based on the sine function to simulate the dynamic changes of traffic states and improve the generalization ability of the model; and combines the GRU gating mechanism to enhance the model's ability to capture feature information.
[0030] The Traveling Salesman Problem (TSP) is a classic combinatorial optimization problem, and its goal is to find the shortest path so that the salesman can visit each given city once and only once and finally return to the starting city. In the present invention, the TSP problem with traffic states will be considered, and its instance is defined as a set of N nodes to be visited, where each node has the following characteristics: Node coordinates: Each node i corresponds to a two-dimensional node coordinate , representing the position of the node in space. These coordinates are used to determine the distance between nodes subsequently.
[0031] Special feature (traffic state): In addition to the coordinate information, by introducing a traffic state as a new feature, this traffic state represents the congestion degree of reaching node i, thus affecting the calculation of the travel time between nodes.
[0032] A solution is composed of a sequence of nodes visited in order. To describe the specific structure of the solution, a position feature is used to represent the position of node i in the solution. That is represents the permutation order of node i in the final solution. For example, if node i is in the 3rd position, then . This representation method can clearly understand the order and structure of each node in the path.
[0033] A complete solution is composed of a sequence and each node is visited and only visited once. By finding an optimal or sub-optimal position permutation, the total travel time is minimized.
[0034] Example 1 This example discloses a vehicle path planning method with traffic status based on dual - collaborative Transformer; As Figure 1 shown, the vehicle path planning method with traffic status based on dual - collaborative Transformer includes: Step S1, construct a set of cities with traffic status, obtain the distances and time costs between cities, construct an objective function with the goal of minimizing the total time cost, and obtain an initial solution for vehicle path planning; Let the set of cities be where is the number of cities. Each city has coordinates .
[0035] Each city has a traffic status , and the traffic status reflects the congestion degree of arriving at city j. At the same time, let represent the speed of arriving at city i.
[0036] The Euclidean distance between city i and city j is:
[0037] Based on this distance, the time cost from city to city is:
[0038] In the formula, is the distance from city i to city j; is the traffic status of arriving at city j; Furthermore, introduce a binary decision variable to describe whether to select the path from city i to city j, and the definition of
[0039] To avoid self - loops (i.e., going from oneself to oneself), for all , there is .
[0040] Construct an objective function to minimize the total time cost, expressed as:
[0041] To ensure the rationality of the model and the feasibility of actual operation, the following constraint conditions are set: Flow balance constraint: Each city is visited only once. For each : , representing the path leaving city ; , representing the path arriving at city ; Subtour elimination (MTZ constraint): To eliminate subtours with less than N nodes, introduce continuous variables , for , satisfying: , , ; where .
[0042] Binary variable constraint: The decision variable is a binary variable, satisfying: , ; Combining the above objective function and constraint conditions, the final mathematical model is as follows:
[0043] ,
[0044] ,
[0045] ,
[0046] ,
[0047] , .
[0048] Step S2, optimize the initial solution of the vehicle path planning using the Markov decision process.
[0049] Among them, the core of the Markov decision process is the automatic selection and real-time adjustment of the strategy. Based on the feedback of the traffic state and the results of path adjustment, the strategy continuously iteratively adjusts the composition of the solution, and finally converges to a high-quality path. To this end, the entire optimization process is modeled as a Markov decision process (MDP), where each decision step corresponds to a state-action pair, and the goal of the model is to maximize the long-term reward (i.e., minimize the total travel time) by selecting appropriate actions at each step.
[0050] The selection of each node pair and the corresponding local operations (such as 2-opt, insertion, or swapping) are carried out based on considering the traffic state. The model continuously adjusts the path to explore better solutions in the solution space.
[0051] In addition, the flowchart of the MDP is as shown Figure 2 to further clearly show how the model makes decisions based on the current state in each round of iteration. Specifically, at each time step, the model selects an action and applies the corresponding operation according to the state of the current solution, the available actions, and the traffic state. After several iterations, the model gradually approaches the optimal solution. During the optimization process, the introduction of the traffic state greatly improves the model's response ability to environmental changes, enabling the final solution to effectively avoid traffic congestion or other adverse factors, obtain a more optimized path, and thus achieve a lower cost.
[0052] Among them, the Markov decision process is as follows: For each instance containing nodes, the state describes the current solution , which contains the features of each node and the location. Among them, the node features are jointly derived from the node coordinates and the traffic state, that is: ; Among them, is the solution in the current state; represents a sequence of solutions; , this state not only reflects the individual information of each node but also the dynamic changes of the entire system, providing comprehensive input features for subsequent action selection; Each action means that is adjusted through paired operators; The reward function is defined as: , ,
[0053] Among them, is the comprehensive reward value; is the gap between the current solution and the historical optimal solution; is the indicator function; is the improvement threshold; a magnitude of positive bias, denotes the indicator function, denotes the improvement threshold, at the current state up to time step is the optimal solution found so far; policy is parameterized using the IDACT model with parameters. At each time step t, the action (i, j) is obtained by sampling from a stochastic policy during training and inference. The policy generates appropriate action decisions by learning the connection between state-action pairs the next state is derived from the current state , and is obtained by adjusting for a given pair of nodes using paired operators.
[0054] Step S3, input the optimized initial solution into the trained dual collaborative Transformer model, and output the optimal path under changing traffic states; The dual collaborative Transformer model enhances the model's ability to handle complex constraints and dynamic traffic states by encoding solutions using embeddings from different aspects. Combining Figure 3 , the IDACT model consists of two main parts: the IDAC encoder and the IDAC decoder. In the IDAC encoder, the model processes information by calculating the self-attention correlation for each aspect separately. Specifically, the original cross-aspect reference attention mechanism is adopted, enabling the model to effectively use the attention information of one aspect as a reference to provide auxiliary information for another aspect. This mechanism helps the information fusion between aspects and improves the model's expression ability in complex environments. In the IDAC decoder, the model collects action assignment suggestions from two aspects and synthesizes them to generate the final decision. In this way, IDACT can achieve more efficient and accurate path optimization under the collaborative action of multi-dimensional features.
[0055] The embeddings for learning will be introduced in detail below, namely the node feature embedding (NFE) for node representation and the position feature embedding (PFE) for position representation.
[0056] Node Feature Embedding (NFE): In the IDACT model, Node Feature Embedding (NFE) is used to represent the feature information of each node. Different from traditional methods, the NFE design not only includes the two-dimensional coordinate information of the node, but also introduces the traffic state feature of the node. Specifically, for each node i, its node feature is initialized as a linear projection with an output dimension of dim = 64. The reason for adding the embedding of traffic state in NFE is that traffic state plays a crucial role in the TSPTS problem. Traffic state can greatly affect the travel time between nodes. Therefore, embedding the traffic state into the node feature allows the model to more comprehensively consider the dual impacts of space and traffic when encoding node information. This design enables the model to jointly model the spatial feature and traffic state feature of the node in subsequent network layers, thereby improving the optimization performance of the model.
[0057] Positional Feature Embedding (PFE): Corresponding to the Node Feature Embedding (NFE), the model follows the Positional Feature Embedding (PFE) of the original DACT to encode the position of the node in the path. To retain the path order information, the cyclic position encoding design based on Cyclic Gray Encoding (CPE) is maintained. Specifically, the positional feature embedding is a real-valued vector, initialized as dim = 64, to represent the order of node i in the path. Although the position encoding itself does not directly contain the information of traffic state, to further enhance the sensitivity of the positional feature to traffic state, the position encoding is adjusted so that it can reflect the impact of traffic state on node position to some extent. The core idea of this adjustment is that although there is no direct interaction between the position encoding and the embedding of traffic state, the model combines the features of the two, enabling the positional feature to be indirectly affected by traffic state during calculation. For example, when the traffic flow is large, the distance between some nodes may be longer, thus affecting the path selection. This combination method enables the model to capture the correlation between node position and traffic state at a higher level, thereby optimizing path planning and time prediction.
[0058] A key mechanism in the IDACT model is the Cross-Attention Mechanism. In the traditional self-attention mechanism, the model only exchanges information through its own features. However, in the IDACT proposed in this embodiment, features of different aspects (such as node features and location features) can reference and influence each other. This cross-aspect attention mechanism enables the attention information of one aspect to effectively serve as a reference for another aspect, helping to better integrate the spatial location and traffic state features of nodes, thereby improving the accuracy of overall path optimization. Specifically, in the IDAC encoder, first, the self-attention of each aspect is calculated, and then through the cross-aspect reference attention mechanism, the attention correlation of one aspect is used as auxiliary information and integrated into the calculation of another aspect. This design not only improves the learning ability of the model but also enhances the model's ability to handle complex traffic environments, especially when the traffic conditions between nodes change significantly.
[0059] Overall, the working process of the IDAC encoder is divided into several key steps to fully process the embedded features and generate accurate output results. Based on the original encoder framework, several adjustments are made to enhance the ability to capture and represent features. The encoder consists of L = 3 stacked IDAC encoding layers, and each encoding layer contains an interactive part of node feature encoding and location feature encoding. In each IDAC encoder, relatively independent encoding streams of node feature embeddings (NFEs) and location feature embeddings (PFEs) are maintained, and these embeddings are processed independently during the encoding process to ensure that the information between different features can be effectively retained and extracted. Among them, the node feature embedding fuses the node coordinates and traffic state. The location feature embedding captures the location information of the nodes. As shown in the following formula, each embedding stream contains a shared two-aspect collaborative attention sublayer (IDAC-Att) and an independent feed-forward network sublayer (FFN). IDAC-Att takes two groups of embeddings as inputs simultaneously and outputs their respective enhanced embeddings, namely and . Each sublayer is followed by a skip connection and layer normalization, which is the same as the original transformer structure. Such a structure helps to alleviate the problem of gradient disappearance and makes the transfer of features smoother.
[0060] Specifically, the update steps of the node feature embedding are as follows: First, the input features are normalized through layer normalization to stabilize the training process, and then enter the feed-forward network (FFN). This layer of network non-linearly transforms the features through an activation function to facilitate the extraction of higher-order feature representations. The update mechanism of the location feature embedding is similar, and its calculation formula is as follows: ; ; In the formula, is the node embedding of the l th layer; is layer normalization; is the node embedding after enhanced embedding; is the feed-forward propagation layer; is the node embedding after forward propagation and layer normalization; is the position embedding after forward propagation and layer normalization; is the position embedding after enhanced embedding; IDAC-Att (bi-aspect collaborative attention): This IDAC-Att sublayer is the core component of the entire encoder. It not only adds the self-attention mechanism in each layer but also uses a feed-forward network (FFN) to enhance the richness of features. Enhance the embedding from the perspective of each embedding set itself, and at the same time utilize the attention correlation from another perspective to achieve collaboration. For two sets of embeddings and , first calculate the self-attention correlation from two aspects respectively:
[0061] Among them, is a trainable parameter matrix, and are the attention correlation scores of the node embedding and the position embedding respectively, and represent the node embedding and the position embedding respectively, is a conventional operation in the attention mechanism.
[0062] Independent matrix is used to calculate the query and the key. Through softmax, the obtained correlation is normalized to and . Note that these correlations are calculated from their respective perspectives, eliminating potential noise, which helps to correctly handle the incompatible node pair relationships in the solution.
[0063] Then, through the cross-aspect reference attention mechanism, the calculated correlations are shared between the two aspects as additional references for comparison and collaboration: ; ; Among them, is a trainable parameter matrix for generating each aspect value, while is a parameter matrix for generating the reference value. Finally, use multi-head attention to obtain and , as follows: ; ; Among them, , representing the corresponding attention head value, DAC-Att refers to the two-way collaborative attention layer, and is a trainable parameter matrix, respectively refer to the enhanced node embedding and position embedding. In the model of this embodiment, , and .
[0064] In the FFN sublayer, the FFN sublayer has only one hidden layer, with 64 hidden units, and uses RELU as the activation function. The parameters of FFN_h and FFN_g are independent for each group of embeddings.
[0065] In summary, through multiple processes of multi-layer embedding, IDAC-Att bidirectional attention layer, cross-reference mechanism, and feed-forward network, the encoder enhances the interaction ability between node features and position features. This design can not only better capture the complex relationships between nodes, but also improve the encoding expression ability, providing a more refined feature input for subsequent entry into the decoder.
[0066] Joint encoding of multi-dimensional input features. The input features consist of two parts: node coordinates , and traffic status . They are concatenated to form a three-dimensional feature vector , which is passed into the subsequent network structure. At the same time, it is used for subsequent path calculation, enabling the model to comprehensively consider the interaction relationship of features during subsequent learning.
[0067] Let the input coordinate matrix be , and the traffic status be . The joint feature encoding is denoted as: ; Its corresponding joint encoding feature matrix is: ; This enables the model to simultaneously consider the influence of position and traffic status, thereby handling more complex environmental factors in path optimization.
[0068] Generally speaking, in the IDAC decoder, the two groups of embeddings obtained after being processed by the encoder and First, it passes through the max - pooling sub - layer and then enters the multi - head compatibility (MHC) sub - layer, which independently generates diverse node - pair selection suggestions from their respective perspectives. These suggestions are then aggregated through the feed - forward aggregation (FFA) sub - layer to finally generate the desired output result.
[0069] Max - pooling: For each set of embeddings, the max - pooling sub - layer is used to process them, aggregating the global representations of all embeddings into each embedding, thus ensuring the effective transmission of information.
[0070] Multi - head compatibility (MHC): The features after max - pooling are further input into the multi - head compatibility layer. The compatibility sub - layer calculates the attention correlation for each pair of embeddings, and the resulting correlation magnitude is , and is regarded as the proposal distribution for node - pair selection. Compatibility is calculated based on multiple heads to ensure diversity. The attention score matrices and are independently calculated from two aspects respectively, and these matrices will be used to represent the compatibility of node - pairs under a specific head.
[0071] These proposal distributions are different, providing a rich pool of proposals for the subsequent FFA layer, making the model more flexible and robust.
[0072] Feed - forward aggregation (FFA): Once all the proposals are collected from two aspects, a four - layer FFN is used, and the ReLU activation function is adopted to aggregate these proposals: ; where m = 4 is the number of multi - heads; the output is a scalar, representing the possibility of selecting the node - pair as an action. Through layer - by - layer aggregation, the IDAC decoder can ensure that the representation of each node fully considers the global information, thus achieving more refined inference. Subsequently, is applied, where is used to control the entropy, and the infeasible node - pairs are masked as . Finally, the possibilities are normalized through the Softmax function to obtain the final action distribution .
[0073] Finally, through the collaborative work of the above - mentioned layers, IDAC ensures that the significant features of the input embeddings are retained, diversity is guaranteed, and finally a higher - quality action distribution is obtained.
[0074] Utilize the dynamic update mechanism to better handle the introduced traffic states: The dynamic traffic states are updated through a time-step strategy with added random perturbations. The traffic states are based on the current time step and the periodic factor for adjustment, while adding random perturbations to simulate uncertainties.
[0075] Initial traffic state: Let be the traffic state vector at time step , and the initial state is .
[0076] Time factor: The periodic change is controlled by the time factor, i.e., .
[0077] Random perturbation: Let , i.e., Gaussian noise with a standard deviation of . Therefore, the traffic state update can be expressed as: ; where is the traffic state at time step , is the time factor, is the random perturbation.
[0078] The design of the time factor is to simulate the periodic fluctuations of traffic states. The sine function has the characteristic of periodicity and is very suitable for simulating real-world traffic conditions, such as the morning and evening rush hours and off-peak hours every day. At the same time, the time factor selects the additive form during the training stage to simulate the gradual accumulation of traffic states (i.e., increasing congestion), while using the subtractive form during the testing stage to reflect the traffic slowdown in reality (such as the smooth traffic after the rush hour). Through this design, the model can be effectively adapted to different environmental conditions.
[0079] To prevent the algorithm from falling into a local optimal solution, a dynamic perturbation strategy is adopted to restore it to the optimal solution. In the function, to judge whether the current reward value is greater than to update the optimal solution. If the quality of the solution has not been improved within consecutive steps, then perturbation is performed and the current solution is restored to the previous best solution.
[0080] Let be the reward at time step , be the state of the solution. The perturbation mechanism is executed according to the following rules: (i) If , it means that a better solution has been found, and then the state is updated.
[0081] (ii) If , then a count is accumulated , and when , a perturbation is triggered and restored to the optimal solution , that is ; The perturbation mechanism is through the solution reset operation (reset part), and at the same time jumps out of the local optimum through the perturbation strategy.
[0082] An improved reinforcement learning method based on Proximal Policy Optimization (PPO) is adopted, and combined with the Curriculum Learning (CL) strategy to improve the sample efficiency. The PPO algorithm is trained through n-step return estimation, and the curriculum learning strategy is used to optimize the training effect, enabling the model to gradually adapt to more complex tasks.
[0083] The core idea of curriculum learning is to gradually increase the difficulty of the training tasks, ensuring that the model can obtain more efficient training effects during the step-by-step learning process from simple to complex. In existing methods, the training steps are usually limited to a small value to reduce the training cost, but this also makes it difficult for the model to see high-quality solutions during training, which may lead to high variance. Especially when the bootstrapping value function estimates future returns, the model may not have enough knowledge to make accurate predictions.
[0084] To overcome the above problems, an improved curriculum learning strategy is introduced. This strategy starts by gradually providing higher-quality initial solutions as the starting point of training, enabling the model to observe and learn relatively good solutions from the beginning, thereby reducing the variance in value function estimation. In addition, as training progresses, the difficulty of the tasks is gradually increased, allowing the model to learn in more challenging scenarios, thus obtaining more stable and efficient training effects.
[0085] In actual operation, the initial state is updated through a gradually improving strategy. These high-quality solutions are obtained through a small number of improvement steps by the current strategy, which can effectively make the training tasks gradually become more complex. The number of initial training steps slowly increases as the number of training epochs increases, enabling the model to better adapt to the learning of high-quality solutions during the continuous improvement process.
[0086] The training algorithm is based on n-step Proximal Policy Optimization (PPO) combined with a curriculum learning strategy, while following the Actor-Critic network structure of PPO. First, the model initializes the state by randomly generating a batch of training instances in each epoch, and then improves these states through the curriculum learning strategy to make them higher-quality initial solutions. During the training process, the model collects experiences and uses n-step return estimation to optimize the model. The objective function of PPO and the clipped value function loss are used for updating, thus ensuring effective improvement of the policy and the value function.
[0087] The curriculum learning strategy gives the model relatively simple tasks at the beginning of training and gradually increases the difficulty. In this way, the model can better adapt to complex tasks, utilize samples more effectively, and improve the overall training efficiency.
[0088] In addition to the basic multi-head attention mechanism and Batch Normalization, a Gated Recurrent Unit (GRU) mechanism is introduced to enhance the model's ability to capture sequential information. The GRU gating mechanism includes two main gates: the update gate and the reset gate.
[0089] GRU is located in the Critic module and can capture the temporal dependencies in the input sequence, enabling the model to retain important historical information. This helps the Critic accurately estimate the state value by modeling dynamic time changes, thereby enhancing the policy optimization ability of the Actor. By selectively updating and retaining features, GRU effectively handles the long-term dependence problem and improves the overall performance and stability of the model. The entire PPO process is as Figure 4 shown.
[0090] To better demonstrate the impact of IDACT on the performance of the Traveling Salesman Problem with Traffic States (TSPTS), it will be shown in the following experimental section. To avoid the differential impact of speed on the results, a fixed speed of 1 is used to eliminate the differences and ensure fairness. Secondly, the experimental scenarios all consider the impact of traffic states. The compared models introduce traffic states to affect the final total travel time, while IDACT adopts the proposed dynamic update mechanism and other innovations to verify the effectiveness of the model. Three different scales of the same instances with N = 20, N = 50, and N = 100 are selected for training and testing. L20-GPUs are used for training and testing, and 1000 out of 10000 instances in the test set are selected for testing. The results obtained are as Figure 5 shown. To verify the effectiveness of the IDACT model in solving the TSPTS problem, in Figure 5IDACT was compared with five existing methods: DACT, AM-sampling, MDAM-BS, FER-POMO, and FER-AM. In all experiments, the results of DACT were selected as the baseline method for comprehensive comparison. All test results were based on the average of 1000 instances, and the final cost, the gap with the baseline model, and the running time in the test were reported respectively.
[0091] Without data augmentation, the performance of IDACT and the baseline method (DACT) on different datasets was first analyzed. For each dataset (TSP20, TSP50, and TSP100), IDACT showed significant performance improvement: in TSP20, the final cost of IDACT increased by 1.83% compared with DACT; in TSP50, the final cost of IDACT increased by 5.29% compared with DACT; in TSP100, IDACT achieved a significant improvement on this dataset, and the final cost decreased by 10.1% compared with DACT.
[0092] These results indicate that the optimization effect of IDACT on the TSPTS problem far exceeds other existing methods. Especially in large-scale instances (such as TSP100), IDACT demonstrated significant performance advantages. This "cliff-like" performance improvement proves the effectiveness of IDACT in dealing with the TSPTS problem, especially when considering the impact of traffic conditions on costs.
[0093] However, although IDACT has made obvious improvements in cost, its running time has increased. When T = 5k, the running time of IDACT is longer than that of DACT. Nevertheless, the performance improvement achieved by trading time is still worthy of recognition. Generally speaking, the performance improvement of IDACT is particularly prominent in TSPTS problems of different scales, showing its potential in practical applications.
[0094] To further verify the robustness of IDACT, experiments on data augmentation were conducted, focusing on comparing the performance differences between IDACT and DACT after data augmentation. Under two data augmentation conditions (2 times augmentation and 4 times augmentation), the performance of IDACT was observed respectively. When augmented 2 times, it performed best: in TSP20, the performance increased by 0.89%; in TSP50, the performance increased by 2.84%; in TSP100, the performance increased by 5.61%.
[0095] These results indicate that with the help of data augmentation, IDACT can still maintain its reliability and effectiveness, and its performance continues to improve. This finding further validates the robustness and adaptability of IDACT. Especially in the face of increasing data diversity, IDACT can effectively improve the generalization ability of the model.
[0096] Embodiment 2 This embodiment discloses a vehicle path planning system with traffic status based on dual collaborative Transformers; As Figure 6 shown, the vehicle path planning system with traffic status based on dual collaborative Transformers includes: An initial solution acquisition module, configured to: construct a set of cities with traffic status, obtain the distances and time costs between cities, construct an objective function with the goal of minimizing the total time cost, and obtain an initial solution for vehicle path planning; An initial solution optimization module, configured to: optimize the initial solution of the vehicle path planning by using a Markov decision process, An optimal path acquisition module, configured to: input the optimized initial solution into a trained dual collaborative Transformer model and output the optimal path under changing traffic status; Among them, the dual collaborative Transformer model represents the feature information of each node through node feature embedding, encodes the position of the node in the path through position feature embedding, and uses a cross-aspect reference attention mechanism to integrate the spatial position and traffic status features of the nodes, and outputs the optimal path under changing traffic status. Embodiment 3 The purpose of this embodiment is to provide a computer-readable storage medium.
[0097] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the vehicle path planning method with traffic status based on dual collaborative Transformers as described in Embodiment 1.
[0098] Embodiment 4 The purpose of this embodiment is to provide an electronic device.
[0099] An electronic device includes a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the vehicle path planning method with traffic status based on dual collaborative Transformers as described in Embodiment 1.
[0100] In the devices of the above Second, Third, and Fourth Embodiments, the steps involved correspond to those of the First Method Embodiment. For the specific implementation, reference may be made to the relevant description part of the First Embodiment. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to execute any method in the present invention.
[0101] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device for execution by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0102] Although the specific implementation of the present invention has been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solution of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.
Claims
1. A vehicle path planning method with traffic states based on a dual-collaborative Transformer, characterized in that, Including: Construct a set of cities with traffic states, obtain the distances and time costs between cities, construct an objective function with the goal of minimizing the total time cost, and generate an initial solution for vehicle route planning; Use the Markov decision process to optimize the initial solution of the vehicle route planning, input the optimized initial solution into the trained dual collaborative Transformer model, and output the optimal route under changing traffic states; Among them, the dual collaborative Transformer model represents the feature information of each node through node feature embedding, encodes the position of the node in the route through position feature embedding, and uses the cross-aspect reference attention mechanism to integrate the spatial position and traffic state features of the node, and outputs the optimal route under changing traffic states.
2. The vehicle path planning method with traffic states based on the dual-collaborative Transformer according to claim 1, characterized in that The process of constructing a set of cities with traffic states, obtaining the distances and time costs between cities, constructing an objective function with the goal of minimizing the total time cost, and obtaining the initial solution of vehicle route planning is as follows: Let the set of cities be , where is the number of cities; each city has coordinates ; each city has a traffic state , and , the traffic state reflects the congestion level of reaching city j. At the same time, let represent the speed of reaching city i; Calculate the Euclidean distance between city i and city j. Based on the Euclidean distance, calculate the time cost from city to city as follows: ; In the formula, is the distance from city to city ; is the traffic condition for arriving at city ; Introduce binary decision variables , which are used to describe whether to select the path from city i to city j, and the definition is as follows: ; Construct an objective function to minimize the total time cost, expressed as: 。 3. The vehicle path planning method with traffic status based on dual collaborative Transformers according to claim 1, characterized in that, The process of using the Markov decision process to optimize the initial solution of the vehicle route planning is as follows: For each instance containing nodes, the state describes the current solution , which contains each node and location features. Among them, the node features are combined by node coordinates and traffic states, that is: , where ; Each action indicates that is adjusted by a pair of operators, and the reward function is defined as: ; ; ; Among them, an amplitude of forward bias, represents an indicator function, represents an improved threshold, is the optimal solution found up to the time step in the current state; the policy is parameterized using the IDACT model with parameters. At each time step t, the action (i, j) is obtained by sampling the stochastic policy during training and inference. The policy generates appropriate action decisions by learning the connection between state-action pairs; the next state is derived from the current state , and is obtained after adjustment by performing pairwise operators on the given node pair .
4. The vehicle path planning method with traffic status based on the dual-collaborative Transformer according to claim 1, characterized in that, The dual collaborative Transformer model includes an IDAC encoder and an IDAC decoder; in the IDAC encoder, the model processes information by calculating the self-attention correlation of each aspect respectively; in the IDAC decoder, the model collects action assignment suggestions from two aspects and synthesizes them to generate the final decision.
5. The vehicle path planning method with traffic status based on the dual collaborative Transformer according to claim 1, characterized in that, The IDAC encoder is composed of stacked IDAC encoding layers, and each encoding layer contains an interactive part of node feature encoding and position feature encoding; in each IDAC encoder, the relatively independent encoding streams of node feature embedding and position feature embedding are maintained.
6. The vehicle path planning method with traffic status based on the dual-collaborative Transformer according to claim 1, wherein, In the IDAC decoder, the two groups of embeddings obtained after being processed by the encoder first pass through the max pooling sublayer and then enter the multi-head compatibility sublayer to independently generate diverse node pair selection suggestions from their respective perspectives. The suggestions are then summarized by the feed-forward aggregation sublayer to finally generate the optimal route.
7. The vehicle path planning method with traffic status based on a dual-collaborative Transformer according to claim 1, wherein, A reinforcement learning method based on proximal policy optimization is also adopted, and a curriculum learning strategy is combined to improve the sample efficiency.
8. A vehicle path planning system with traffic status based on dual collaborative Transformers, characterized by: Including: An initial solution acquisition module, configured to: construct a set of cities with traffic states, obtain the distances and time costs between cities, construct an objective function with the goal of minimizing the total time cost, and obtain the initial solution of vehicle route planning; An initial solution optimization module, configured to: use the Markov decision process to optimize the initial solution of the vehicle route planning, An optimal route acquisition module, configured to: input the optimized initial solution into the trained dual collaborative Transformer model, and output the optimal route under changing traffic states; Among them, the dual collaborative Transformer model represents the feature information of each node through node feature embedding, encodes the position of the node in the route through position feature embedding, and uses the cross-aspect reference attention mechanism to integrate the spatial position and traffic state features of the node, and outputs the optimal route under changing traffic states.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the vehicle path planning method with traffic status based on dual collaborative Transformers as described in any one of claims 1-7.
10. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the vehicle path planning method with traffic status based on dual collaborative Transformers as described in any one of claims 1-7.