Road-vehicle cooperative path planning method based on traffic flow prediction

By integrating radar data and traffic flow prediction, Markov decision-making model and D3QN model are established, and the path planning problem of autonomous vehicles at complex intersections is solved, and accurate path planning and safe driving are achieved.

CN120299261AInactive Publication Date: 2025-07-11山西省交通科技研发有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510786979.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The path planning of autonomous vehicles at complex traffic intersections faces challenges in environmental uncertainty, perception system limitations, real-time requirements, decision-making complexity and multi-data fusion, which makes it difficult for sensors to fully and accurately perceive the environment and make it difficult to achieve accurate path planning.

Method used

Through roadside units, the radar point cloud data is integrated, the feature extraction and fusion is used to form a road traffic flow map, combined with ST-Conv for prediction, a Markov decision model (MDP) agent is established, and the decision planning is used to use the D3QN double-layer Q-net model to realize the adaptive decision-making and path planning of the agent.

Benefits of technology

It realizes accurate perception, traffic flow prediction and rapid decision-making of autonomous vehicles at complex intersections, and improves driving safety and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299261A_ABST
    Figure CN120299261A_ABST
Patent Text Reader

Abstract

A road-vehicle cooperative path planning method based on traffic flow prediction comprises the following steps: step 1, a road side unit integrates collected radar point cloud data and data uploaded by a vehicle to form a road traffic flow map, performs road traffic flow prediction and outputs a traffic congestion map; 2, establishing a Markov decision model (MDP) proxy for each automatic driving vehicle, and initializing an MDP proxy model according to the traffic congestion map; step 3, performing decision planning by using a D3QN double-layer Q-net model, and controlling a decision strategy by modifying parameters of an MDP agent by a D3QN network to realize agent adaptive decision; and meanwhile, the D3QN model carries out path planning according to the decision result, and outputs the planned path to the automatic driving vehicle, so that multi-vehicle cooperative driving is realized. According to the method, the environment can be accurately sensed, the traffic flow change can be predicted, the decision can be quickly made, and effective communication with traffic infrastructures can be realized, so that the driving safety and efficiency of an automatic driving vehicle at a complex intersection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of path planning, and specifically relates to a vehicle-road collaborative path planning method based on traffic flow prediction. Background Art

[0002] At present, autonomous driving technology is becoming increasingly perfect, and its potential in improving road safety, optimizing traffic flow, reducing environmental pollution, etc. has been gradually recognized. However, the path planning of autonomous vehicles at complex traffic intersections is still a technical problem. This challenge mainly comes from the uncertainty of the environment, the limitations of the perception system, the requirements of real-time performance, the complexity of decision-making, the challenges of multi-source data fusion, the practical application difficulty of vehicle-road collaborative technology, and the computational complexity of algorithms. Summary of the Invention

[0003] The present invention provides a vehicle-road collaborative path planning method based on traffic flow prediction to solve the defects in the prior art that when an autonomous vehicle is in a complex traffic section, the accuracy of its radar and other sensing devices is limited and cannot meet the requirements of positioning and planning. Secondly, the obstruction of vehicles and other obstacles makes it impossible for the radar and other sensors to comprehensively and accurately sense all objects on the road, and it is difficult for autonomous vehicles to complete accurate and reasonable path planning.

[0004] The present invention is achieved through the following technical solutions: A vehicle-road collaborative path planning method based on traffic flow prediction includes the following steps: Step 1: The roadside unit integrates the collected radar point cloud data and the data uploaded by vehicles, then uses CNN for local feature extraction (LFE) and fuses the features to form a road traffic flow map; secondly, the map is input into ST-Conv for road traffic flow prediction, and a traffic congestion map is output; Step 2: A Markov decision model (MDP) agent is established for each autonomous vehicle. The Markov decision model (MDP) agent is represented by a quadruple, and the variables in the quadruple are the state space, action space, state transition probability, and reward function respectively. Finally, the MDP agent model is initialized according to the traffic congestion map; Step 3: The D3QN double-layer Q-net model is used for decision-making and planning. The D3QN network indirectly controls the decision-making strategy by modifying the four parameters of the MDP agent to achieve intelligent agent adaptive decision-making; at the same time, the D3QN model will perform path planning according to the decision-making result and output the planned path to the autonomous vehicle to achieve multi-vehicle collaborative driving.

[0005] For the vehicle-road collaborative path planning method based on traffic flow prediction as described above, the road traffic flow prediction method in Step 1 includes the following steps: Step 1: The prediction module uses a CNN-based local feature extraction method for feature extraction. The input traffic flow raster map is divided into multiple sectors, and each sector is represented by a two-channel raster map. Step 2: The prediction module uses TCN for time feature extraction and uses the gated linear unit (GLU) to connect features at different time steps to transfer residual information between stacked layers of the network. Step 3: Use a multi-layer perceptron (MLP) to map the learned single-step prediction features onto the raster map to form a road grid traffic flow map with aligned spatio-temporal dimensions.

[0006] A vehicle-road collaborative path planning method based on traffic flow prediction as described above. The specific operation of representing each sector by a two-channel raster map in Step 1 is as follows: The first channel is the static layer of static obstacles in the environment, and the second channel is the dynamic layer of dynamic obstacles in the environment, that is, moving vehicles. The shape of its internal information is , and then use CNN to obtain the vector Secondly, use a graph convolutional network to consider the mutual influence between sectors. The specific expression is: where σ is the activation function, is the adjacency matrix, is the corresponding degree matrix, is the weight matrix, is the input and output.

[0007] A vehicle-road collaborative path planning method based on traffic flow prediction as described above. In Step 2, the state space is in the vehicle's driving trajectory. The evolution of the state space depends on the change of the occupied road section. Therefore, the definition of the state space is closely related to the current road section. The specific definition elements are , where represents the current road, represents the congestion level of the road connected to (unconnected roads are regarded as infinitely congested), represents the length of the road related to (unconnected roads are regarded as infinitely long), and d represents the destination of the current journey.

[0008] A vehicle-road collaborative path planning method based on traffic flow prediction as described above. In Step 2, the value of the action space is based on the number of connected roads. The vehicle's action is the same as the number of roads connected to the current road section. Different numerical outputs represent the roads that the vehicle intends to navigate to next.

[0009] A vehicle-road collaborative path planning method based on traffic flow prediction as described above, where the state transition probability refers to the probability of transitioning from the current state to the next state . This probability depends on the action output by the vehicle and is calculated through , where is the action.

[0010] A vehicle-road collaborative path planning method based on traffic flow prediction as described above. As the vehicle iteratively optimizes its path planning strategy, the output action at each step plays a key role in reward calculation. Selecting road segments with shorter distances, lower congestion, and closer to the destination will increase the reward positively.

[0011] A vehicle-road collaborative path planning method based on traffic flow prediction as described above. The formula for reward calculation is as follows: , where the parameters and are used to adjust the reward, represents the reward penalty, which is positive when the subsequent road is closer to the destination and negative when it is farther.

[0012] A vehicle-road collaborative path planning method based on traffic flow prediction as described above. The specific operations of the decision-making and planning in step three include the following steps: Step (1): Provide the obtained environmental information and the predicted traffic flow information to the MDP intelligent agent. The agent first initializes the network parameters and the experience replay memory. Then, the evaluation network of D3QN outputs an action at each time step, and this action corresponds to the next intersection that the agent intends to navigate to. This action is determined based on the agent's value evaluation (Q) of different actions in the current state; Step (2): After the decision is made, the vehicle executes the action, that is, navigates to the corresponding road, updates the reward parameters according to the reward mechanism. After the action is executed, the corresponding environment will change, and the state becomes ; Step (3): Finally, store the experience tuple composed of the of the current state, the action, the reward, and the next state into the experience replay memory; Step (4): Sample extraction and target value calculation. When the experience replay memory reaches the preset capacity, randomly extract training samples from the memory. For each sample, calculate the target value, and its calculation formula is: , where is the reward of the current state, is the discount factor, represents in the next state Any action that can be taken is a parameter of the target network; Step (5): The training objective of the Critic network is to obtain an approximation of the state-action value function, making the Q value it outputs as close as possible to the actual optimal Q value. By calculating the loss function, its expression is: and use the gradient descent method to update the network parameters whose expression is: ← where, is the learning rate. By continuously repeating the above steps, the MDP agent learns the optimal path planning strategy during the interaction with the environment and can select the most suitable path according to different road states and environmental conditions.

[0013] The advantages of the present invention are: The present invention can accurately perceive the environment, predict traffic flow changes, make decisions quickly, and can communicate effectively with traffic infrastructure, thereby improving the driving safety and efficiency of autonomous vehicles at complex intersections. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0015] Figure 1 is the flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0017] As Figure 1 shown, a road-vehicle collaborative path planning method based on traffic flow prediction includes the following steps: Step 1: The roadside unit integrates the collected radar point cloud data and the data uploaded by the vehicle, then uses CNN for local feature extraction (LFE) and fuses the features to form a road traffic flow map; secondly, the map is input into ST-Conv for road traffic flow prediction, and a traffic congestion map is output; Step 2: Establish a Markov decision model (MDP) agent for each autonomous vehicle. The MDP agent is represented by a quadruple, and the variables in the quadruple are the state space, action space, state transition probability, and reward function. Finally, initialize the MDP agent model according to the traffic congestion map; Step 3: Use the D3QN double-layer Q-net model for decision-making and planning. The D3QN network indirectly controls the decision-making strategy by modifying the four parameters of the MDP agent to achieve adaptive decision-making of the agent. At the same time, the D3QN model will perform path planning according to the decision-making results and output the planned path to the autonomous vehicle to achieve multi-vehicle collaborative driving.

[0018] Specifically, the road traffic flow prediction method in Step 1 of this embodiment includes the following steps: Step 1: The prediction module uses a local feature extraction method based on CNN for feature extraction. The input traffic flow raster map is divided into multiple sectors, and each sector is represented by a two-channel raster map; Step 2: The prediction module uses TCN for time feature extraction and uses the gated linear unit GLU to connect the features of different time steps to transfer residual information between the stacked layers of the network; Step 3: Use a multi-layer perceptron (MLP) to map the learned single-step prediction features to the raster map to form a road grid traffic flow map with aligned spatio-temporal dimensions.

[0019] Specifically, the specific operation of representing each sector by the two-channel raster map in Step 1 of this embodiment is as follows: The first channel is the static layer of static obstacles in the environment, and the second channel is the dynamic layer of dynamic obstacles in the environment, that is, moving vehicles. The shape of its internal information is , and then use CNN to obtain the vector Secondly, use a graph convolutional network to consider the mutual influence between sectors. The specific expression is: where σ is the activation function, is the adjacency matrix, is the corresponding degree matrix, is the weight matrix, is the input and output.

[0020] More specifically, in Step 2 of this embodiment, in the vehicle driving trajectory of the state space, the evolution of the state space depends on the change of the occupied road section. Therefore, the definition of the state space is closely related to the current road section; the specific definition elements are , where represents the current road, Indicates the congestion level of the road connected to (roads that are not connected are considered to have infinite congestion), Indicates the length of the road related to (roads that are not connected are considered to have infinite length), and d represents the destination of the current trip.

[0021] More specifically, the value of the action space in step 2 of this embodiment is based on the number of connected roads. The action of the vehicle is consistent with the number of roads connected to the current road segment. Different numerical outputs represent the roads that the vehicle intends to navigate to next.

[0022] Even more specifically, the state transition probability described in this embodiment refers to the probability of transitioning from the current state to the next state . This probability depends on the action output by the vehicle and is calculated through , where

[0023] is the action.

[0024] Furthermore, the reward function in this embodiment changes as the vehicle iteratively optimizes its path planning strategy. The output action at each step plays a key role in the reward calculation. Selecting road segments with shorter distances, lower congestion, and closer to the destination will increase the reward positively. Moreover, the formula for calculating the reward in this embodiment is: and are used to adjust the reward. represents the reward penalty, which is positive when the subsequent road is closer to the destination and negative when it is farther.

[0025] Even further, the specific operations of the decision-making and planning in step 3 of this embodiment include the following steps: Step (1): Provide the obtained environmental information and the predicted traffic flow information to the MDP intelligent agent. The agent first initializes the network parameters and the experience replay memory. Then, the evaluation network of D3QN outputs an action at each time step. This action corresponds to the next intersection that the intelligent agent intends to navigate to and is determined based on the agent's value evaluation (Q) of different actions in the current state. Step (2): After the decision is made, the vehicle executes the action, that is, navigates to the corresponding road, updates the reward parameter according to the reward mechanism. After executing the action, the corresponding environment will change, and the state becomes . Step (3): Finally, the of the current state, the action, the reward, and the next state The composed experience tuples are stored in the experience replay memory; Step (4): Sample extraction and target value calculation. When the experience replay memory reaches the preset capacity, training samples are randomly extracted from the memory. For each sample, the target value is calculated, and its calculation formula is: , where is the reward for the current state, is the discount factor, represents any action that can be taken in the next state , are the parameters of the target network; Step (5): The training objective of the Critic network is to obtain an approximation of the state-action value function, making the Q value it outputs as close as possible to the actual optimal Q value. By calculating the loss function, its expression is: , and the gradient descent method is used to update the network parameters , and its expression is: ← , where is the learning rate. By continuously repeating the above steps, the MDP agent learns the optimal path planning strategy during the interaction with the environment and can select the most appropriate path according to different road states and environmental conditions.

[0026] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A vehicle-road collaborative path planning method based on traffic flow prediction, characterized in that: It includes the following steps: Step 1: The roadside unit integrates the collected radar point cloud data and the data uploaded by vehicles, then uses CNN for local feature extraction and fuses the features to form a road traffic flow map; secondly, the map is input into ST-Conv for road traffic flow prediction, and a traffic congestion map is output; Step 2: A Markov decision model agent is established for each autonomous vehicle. The Markov decision model (MDP) agent is represented by a quadruple, and the variables in the quadruple are state space, action space, state transition probability, and reward function. Finally, the MDP agent model is initialized according to the traffic congestion map; Step 3: The D3QN double-layer Q-net model is used for decision-making and planning. The D3QN network indirectly controls the decision-making strategy by modifying the four parameters of the MDP agent to achieve intelligent agent adaptive decision-making; at the same time, the D3QN model will perform path planning according to the decision-making result and output the planned path to the autonomous vehicle to achieve multi-vehicle collaborative driving.

2. The vehicle-road collaborative path planning method based on traffic flow prediction according to claim 1, characterized in that: The road traffic flow prediction method in Step 1 includes the following steps: Step 1: The prediction module uses a local feature extraction method based on CNN for feature extraction, divides the input traffic flow raster map into multiple sectors, and represents each sector with a two-channel raster map; Step 2: The prediction module uses TCN for time feature extraction and uses the gated linear unit GLU to connect the features of different time steps to transfer residual information between the stacked layers of the network; Step 3: The multi-layer perceptron is used to map the learned single-step prediction features to the raster map to form a road grid traffic flow map with spatio-temporal dimension alignment.

3. The vehicle-road collaborative path planning method based on traffic flow prediction according to claim 2, wherein: The specific operation of the raster images of the two channels in the above-mentioned step 1 for each sector is as follows: the first channel is the static layer of static obstacles in the environment, and the second channel is the dynamic layer of dynamic obstacles in the environment, that is, moving vehicles. The shape of the internal information The shape of , and then a vector is obtained using CNN Secondly, the graph convolutional network is used to consider the mutual influence between sectors. The specific expression is as follows: where σ is the activation function, is the adjacency matrix, is the corresponding degree matrix, is the weight matrix, is the input and output.

4. A vehicle-road collaborative path planning method based on traffic flow prediction according to claim 1, characterized in that: In the second step described above, in the vehicle driving trajectory, the evolution of the state space depends on the change of the occupied road section. Therefore, the definition of the state space is closely related to the current road section. The specific definition elements are , where represents the current road, represents the congestion level of the road connected to (if not connected, it is regarded as infinitely congested), represents the length of the road related to (if not connected, it is regarded as infinitely long), and d represents the destination of the current trip.

5. A vehicle-road collaborative path planning method based on traffic flow prediction according to claim 1, characterized in that: The value of the action space in Step 2 is based on the number of connected roads. The action of the vehicle is the same as the number of roads connected to the current road section. Different numerical outputs represent the roads that the vehicle intends to navigate to next.

6. The vehicle-road collaborative path planning method based on traffic flow prediction according to claim 1, wherein: The described state transition probability refers to the probability of transitioning from the current state to the next state . This probability depends on the action output by the vehicle and is calculated through , where is the action.

7. A vehicle-road collaborative path planning method based on traffic flow prediction according to claim 1, characterized in that: The reward function optimizes its path planning strategy with vehicle iteration. The output action of each step plays a key role in the reward calculation. Selecting road segments with shorter distances, lower congestion, and closer to the destination will increase the reward positively.

8. A vehicle-road collaborative path planning method based on traffic flow prediction according to claim 7, characterized in that: The formula for calculating the reward is as follows: , where the parameters and are used to adjust the reward, represents the reward penalty, which is positive when the subsequent road is closer to the destination and negative when it is farther.

9. The vehicle-road collaborative path planning method based on traffic flow prediction according to claim 1, wherein: The specific operation of the decision-making and planning in Step 3 includes the following steps: Step (1): Provide the obtained environmental information and the predicted traffic flow information to the MDP intelligent agent; the agent first initializes the network parameters and the experience replay memory. Then, the evaluation network of D3QN will output an action at each time step, and this action corresponds to the next intersection that the agent intends to navigate to. This action is determined based on the agent's evaluation of the value (Q) of different actions in the current state; Step (2): After the decision is made, the vehicle performs an action, that is, navigates to the corresponding road, updates the reward parameter according to the reward mechanism. After the action is completed, the corresponding environment will change, and the state becomes ; Step (3): Finally, store the experience tuple composed of the , action, reward, and the next state in the experience replay memory; Step (4): Sample extraction and target value calculation. When the experience replay memory reaches the preset capacity, randomly extract training samples from the memory. For each sample, calculate the target value, and its calculation formula is: , where is the reward of the current state, is the discount factor, represents any action that can be taken in the next state , are the parameters of the target network; Step (5): The training objective of the Critic network is to obtain an approximation of the state-action value function, making the Q-value it outputs as close as possible to the actual optimal Q-value. By calculating the loss function, its expression is: , and use the gradient descent method to update the network parameters , and its expression is: ← , where is the learning rate. By continuously repeating the above steps, the MDP agent learns the optimal path planning strategy during the interaction with the environment and can select the most suitable path according to different road states and environmental conditions.

Citation Information

Cited By

  • Vehicle-road cooperation path planning decision-making system based on deep learning

    CN120580877A

  • Road train lane line detection method based on image segmentation

    CN121884300A