Automatic driving path planning method and system, electronic equipment and medium
By building a joint action space and collaborative training agent, optimizing signal phase and vehicle paths, the problem of vehicle speed drop in autonomous driving under complex environments and weather conditions is solved, and the traffic efficiency is improved.
Patent Information
- Application Number
- CN202510450980.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Under complex road conditions and extreme weather conditions, the vehicle trajectory prediction error is large, resulting in a decrease in vehicle speed and low traffic efficiency.
By acquiring road traffic data in a specific area, using the space-time coupling matrix to construct a joint action space between the signal control agent and the path planning agent, combined with the Actor-Critic framework and the improved MADDPG algorithm for collaborative training, optimize the signal phase switching time and vehicle driving path, form a hybrid action space, and control traffic lights and vehicles.
Under complex environments and weather conditions, ensure control stability, improve the traffic efficiency of vehicles in specific areas, reduce interference from external factors, and achieve efficient traffic flow.
Smart Images

Figure CN120293170A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of intelligent transportation systems and autonomous driving, and particularly relates to a method, a system, an electronic device and a medium for autonomous driving path planning. Background Art
[0002] Most traffic accidents are caused by human errors, such as fatigue driving, distraction, drunk driving, etc. The autonomous driving system can more reliably comply with traffic rules through sensors, artificial intelligence and precise control algorithms, reducing the possibility of accidents. Based on this, the autonomous driving technology has been gradually proposed.
[0003] However, although the current autonomous driving technology is constantly improving, under complex road conditions and extreme weather conditions, the prediction error of the Long Short - Term Memory (LSTM) prediction model for vehicle trajectory prediction reaches 1.5 - 2.3 meters. In addition, when the simulation environment is deployed to the real scenario, due to differences in complex environments and weather conditions, etc., the effect of the model will decay, resulting in a decrease in vehicle speed, and further bringing the problem of reduced traffic efficiency. Summary of the Invention
[0004] In order to overcome the defect that the effect of the above - mentioned autonomous driving path planning will decay under complex environments and weather conditions, resulting in a decrease in vehicle speed and further bringing low traffic efficiency, the present invention provides an autonomous driving path planning method, including:
[0005] Obtain the road traffic data of a specific area;
[0006] Based on the road traffic data and combined with the spatio - temporal coupling matrix, solve the joint action space of the signal control agent and the path planning agent to obtain the optimal joint action; the optimal joint action includes the optimal signal phase switching time and vehicle driving path within the specific area;
[0007] Use the optimal signal phase switching time and vehicle driving path to control the traffic lights and vehicles within the specific area;
[0008] The joint action space is formed by taking the Cartesian product of the signal control agent and the path planning agent to form a hybrid action space, and then constructing the hybrid action space in combination with the Actor - Critic framework.
[0009] Furthermore, the construction process of the joint action space includes:
[0010] Generate discrete signal control agents according to the green - light duration and switching transition time in traffic lights;
[0011] Generate a continuous path planning agent based on the steering angle and acceleration during vehicle driving;
[0012] Form a hybrid action space by taking the Cartesian product of the signal control agent and the path planning agent;
[0013] Define the joint action in the hybrid action space as the output of the Actor network in the Actor-Critic framework;
[0014] Define the road traffic state and the joint action as the input of the Critic network in the Actor-Critic framework, and locate the joint action taken in the road traffic state as the output of the Critic network;
[0015] Construct a joint action space according to the defined Actor-Critic framework.
[0016] Further, the pre-training process of the joint action space includes:
[0017] Take the road traffic state in the joint action space to be trained as the state quantity of the autonomous driving path planning scenario;
[0018] Take the joint action in the hybrid action space as the action quantity of the autonomous driving path planning scenario;
[0019] Based on the state quantity and the action quantity, use the improved Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm to perform collaborative iterative training on the joint action space to be trained, and obtain the trained joint action space.
[0020] Further, the process of using the improved MADDPG algorithm to perform collaborative iterative training on the joint action space to be trained based on the state quantity and the action quantity, and obtaining the trained joint action space includes:
[0021] Based on the improved MADDPG algorithm, initialize the parameters in the joint action space to be trained under the constraint of the spatio-temporal coupling matrix; the parameters include the network parameters of the Actor-Critic framework;
[0022] Based on the state quantity, use the Actor-Critic framework with initialized parameters to generate the corresponding action quantity, and calculate the value estimation value when the signal control agent and the path planning agent jointly execute the action quantity;
[0023] Iteratively optimize the network parameters of the Critic network and the Actor network in the Actor-Critic framework based on the value estimate to obtain a trained joint action space.
[0024] Further, the iteratively optimizing the network parameters of the Critic network and the Actor network in the Actor-Critic framework based on the value estimate to obtain a trained joint action space includes:
[0025] Based on the value estimate, calculate a loss value using the mean squared error loss function; optimize the network parameters of the Critic network based on the loss value;
[0026] Use road simulation technology to simulate the joint execution of the action quantity by the signal control agent and the path planning agent to obtain simulated traffic data;
[0027] Based on the simulated traffic data, calculate the reward values of the signal control agent and the path planning agent using the reward function; optimize the network parameters of the Actor network based on the reward values;
[0028] Continue to generate the corresponding action quantity based on the Critic network and the Actor network after optimizing the network parameters until the iteration termination condition is met, to obtain a trained joint action space.
[0029] Further, the reward function satisfies the following formula:
[0030]
[0031] where R total is the reward value, α is a preset weight, X is the overall traffic efficiency of the specific area, N is the number of vehicles in the specific area, x i is the actual driving time of the i-th vehicle in the driving data, and is the expected driving time of the i-th vehicle in the driving data.
[0032] Further, before using the optimal signal phase switching time and vehicle driving path to control the traffic lights and vehicles in the specific area, it further includes:
[0033] Based on the optimal signal phase switching time and vehicle driving path in the optimal joint action, use a conflict detection function to determine the traffic conflict detection result;
[0034] If the traffic conflict detection result is that there is a potential traffic conflict, then use a preset correction mechanism to correct the optimal joint action.
[0035] On the other hand, the present invention also provides an autonomous driving path planning system, including:
[0036] A traffic data collection module for obtaining road traffic data in a specific area;
[0037] A dynamic environment modeling module for solving the joint action space of a signal control agent and a path planning agent based on the road traffic data combined with a spatio-temporal coupling matrix to obtain an optimal joint action; the optimal joint action includes the optimal signal phase switching time and vehicle driving path in the specific area;
[0038] An intelligent decision-making module for controlling traffic lights and vehicles in the specific area by using the optimal signal phase switching time and vehicle driving path; the joint action space is formed by taking the Cartesian product of the signal control agent and the path planning agent to form a hybrid action space, and then constructing the hybrid action space by combining it with the Actor-Critic framework.
[0039] Furthermore, the dynamic environment modeling module is also used to construct the joint action space in the following manner:
[0040] Generating discrete signal control agents according to the green light duration and switching transition time in traffic lights;
[0041] Generating continuous path planning agents according to the steering angle and acceleration during vehicle driving;
[0042] Forming a hybrid action space by taking the Cartesian product of the signal control agent and the path planning agent;
[0043] Defining the joint action in the hybrid action space as the output of the Actor network in the Actor-Critic framework;
[0044] Defining the road traffic state and the joint action as the input of the Critic network in the Actor-Critic framework, and positioning the taking of the joint action in the road traffic state as the output of the Critic network;
[0045] Constructing a joint action space according to the defined Actor-Critic framework.
[0046] Furthermore, it also includes a multi-objective optimization module for training the joint action space in the following manner:
[0047] Regarding the road traffic state in the joint action space to be trained as the state quantity of the autonomous driving path planning scenario;
[0048] Take the joint action in the mixed action space as the action quantity for the autonomous driving path planning scenario;
[0049] Based on the state quantity and the action quantity, use the improved multi-agent deep deterministic policy gradient (MADDPG) algorithm to perform collaborative iterative training on the joint action space to be trained, and obtain a trained joint action space.
[0050] Furthermore, the multi-objective optimization module is specifically used to initialize the parameters in the joint action space to be trained based on the improved MADDPG algorithm under the constraint of the spatio-temporal coupling matrix; the parameters include the network parameters of the Actor-Critic framework;
[0051] Based on the state quantity, use the Actor-Critic framework with initialized parameters to generate corresponding action quantities, and calculate the value estimation value when the signal control agent and the path planning agent jointly execute the action quantities;
[0052] Iteratively optimize the network parameters of the Critic network and the Actor network in the Actor-Critic framework based on the value estimation value to obtain a trained joint action space.
[0053] Furthermore, the multi-objective optimization module is specifically used to calculate the loss value based on the value estimation value using the mean squared error loss function; optimize the network parameters of the Critic network based on the loss value;
[0054] Use road simulation technology to simulate the joint execution of the action quantities by the signal control agent and the path planning agent to obtain simulated traffic data;
[0055] Based on the simulated traffic data, calculate the reward values of the signal control agent and the path planning agent using the reward function; optimize the network parameters of the Actor network based on the reward values;
[0056] Continue to generate corresponding action quantities based on the Critic network and the Actor network with optimized network parameters until the iteration termination condition is met, and obtain a trained joint action space.
[0057] Furthermore, the intelligent decision-making module is also used to determine the traffic conflict detection result using the conflict detection function based on the best signal phase switching time and vehicle driving path in the best joint action;
[0058] If the traffic conflict detection result indicates the existence of potential traffic conflicts, use a preset correction mechanism to correct the best joint action.
[0059] On the other hand, the present invention also provides a computer device, which is characterized by comprising: one or more processors;
[0060] The processor is used to store one or more programs;
[0061] When the one or more programs are executed by the one or more processors, the autonomous driving path planning method described in any one of the above is implemented.
[0062] On the other hand, the present invention also provides a computer-readable storage medium, which is characterized in that a computer program is stored thereon, and when the computer program is executed, the autonomous driving path planning method described in any one of the above is implemented.
[0063] Compared with the prior art, the beneficial effects of the present invention are:
[0064] The present invention provides an autonomous driving path planning method, system, electronic device and medium. The method includes: the electronic device acquires road traffic data of a specific area, and based on the road traffic data, uses a spatio-temporal coupling matrix to obtain an optimal joint action. The optimal joint action includes the optimal signal phase switching time and the vehicle driving path, and is used to control traffic lights and vehicles in the specific area. The joint action space is formed by taking the Cartesian product of a signal control agent and a path planning agent to form a hybrid action space, and then is constructed by combining the hybrid action space with the Actor-Critic framework. This method not only makes the decision-making process completely based on objective data, is not interfered by external factors such as weather, ensures the stability of control, but also can improve the traffic efficiency of vehicles in a specific area. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 is a schematic flow chart of the autonomous driving path planning method of the present invention;
[0066] Figure 2 is a schematic structural diagram of the autonomous driving path planning system of the present invention;
[0067] Figure 3 is a schematic structural diagram of the electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] The following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings.
[0069] Embodiment 1:
[0070] An autonomous driving path planning method provided by the present invention, the schematic flow chart is as Figure 1 shown, and includes:
[0071] Step 101: Obtain the road traffic data of a specific area;
[0072] Step 102: Based on the road traffic data and combined with the spatio-temporal coupling matrix, solve the joint action space of the signal control agent and the path planning agent to obtain the optimal joint action; the optimal joint action includes the optimal signal phase switching time and vehicle driving path within the specific area;
[0073] Step 103: Use the optimal signal phase switching time and vehicle driving path to control the traffic lights and vehicles within the specific area; the joint action space is formed by taking the Cartesian product of the signal control agent and the path planning agent to form a hybrid action space, and then constructing it by combining with the Actor-Critic framework.
[0074] An automatic driving path planning method provided by an embodiment of the present invention is applied to an electronic device, and the electronic device can be an intelligent device such as a personal computer (PC), a server, etc.
[0075] In order to accurately and effectively perform automatic driving path planning, in the embodiment of the present invention, a spatio-temporal coupling matrix is used to obtain the optimal joint action, and the traffic lights and vehicles within the specific area are controlled by using the optimal signal phase switching time and vehicle driving path included in the optimal joint action. This method not only makes the decision-making process completely based on objective data, is not interfered by external factors such as weather, ensures the stability of control, but also can improve the traffic efficiency of vehicles in the specific area.
[0076] In order to implement automatic driving path planning, the electronic device can first obtain the road traffic data of a specific area. Specifically, multi-dimensional road traffic data including vehicle position, speed, traffic signal status, and road congestion index can be obtained in real time through roadside sensors, in-vehicle devices, and traffic monitoring systems.
[0077] In order to accurately and effectively perform path planning, the electronic device can establish a data preprocessing pipeline to perform operations such as timestamp alignment, noise filtering, and format standardization on the obtained road traffic data.
[0078] Based on the multi-source fusion road traffic data, the electronic device constructs a spatio-temporal coupling matrix to characterize the spatio-temporal evolution characteristics of the traffic flow. The spatio-temporal coupling matrix establishes a collaborative optimization model between signal control and vehicle paths by quantitatively analyzing the traffic state correlation of each road section in the specific area at different times.
[0079] Specifically, the spatio-temporal coupling matrix describes the interaction of traffic elements from two dimensions of time and space: the time dimension reflects the timing characteristics such as signal cycle and phase difference, and the space dimension characterizes the topological features such as road section passing capacity and turning relationship. During the solution process, the electronic device can adopt a deep reinforcement learning algorithm, model the signal control agent and the path planning agent as collaborative decision-making entities, and finally output the optimal joint action through iterative optimization of the joint action space. The optimal joint action includes the signal phase switching time and the vehicle driving path within the specific area.
[0080] In the embodiment of the present invention, the electronic device adopts a collaborative optimization mechanism based on the Actor-Critic framework to construct a joint action space for signal control and path planning. Specifically, two independent but closely related action domains are defined: the signal control agent is responsible for the discrete action space to optimize the phase switching strategy of traffic lights; the path planning agent is responsible for the continuous action space to dynamically adjust the driving trajectory of the vehicle. Through the Cartesian product operation, the discrete and continuous action spaces are deeply fused to form a unified hybrid action space. Based on the Actor-Critic framework, the collaborative learning and decision-making of the two agents are realized: the Critic evaluates the global value of the joint action and provides the optimization direction for the Actor; the Actor optimizes the signal control strategy and the path planning strategy respectively. This design allows the electronic device to consider both the changes in traffic lights and the adjustment of vehicle driving trajectories at the same time.
[0081] After obtaining the optimal joint action, the electronic device can control the traffic lights and vehicles within the specific area by using the optimal signal switching time and vehicle driving path included in the optimal joint action.
[0082] In the embodiment of the present invention, in order to ensure the smooth operation of the traffic system, a spatio-temporal coupling matrix is introduced to coordinate the signal phase switching time window and the vehicle expected arrival time. When selecting actions, considering the design of the joint action space, the signal control agent and the path planning agent need to work together to jointly determine the optimal joint action. At this time, the spatio-temporal coupling matrix helps to determine whether these actions will cause potential conflicts and guides the signal control agent and the path planning agent to make necessary adjustments.
[0083] In the embodiment of the present invention, it is equivalent to realizing autonomous driving path planning based on a multi-agent reinforcement learning algorithm.
[0084] In order to accurately and effectively construct the joint action space, on the basis of the above embodiments, in the embodiment of the present invention, the construction process of the joint action space includes:
[0085] Generate a discrete signal control agent according to the green light duration and switching transition time in the traffic lights;
[0086] Generate a continuous path planning agent based on the steering angle and acceleration during vehicle driving;
[0087] Form a hybrid action space by taking the Cartesian product of the signal control agent and the path planning agent;
[0088] Define the joint action in the hybrid action space as the output of the Actor network in the Actor-Critic framework;
[0089] Define the road traffic state and the joint action as the input of the Critic network in the Actor-Critic framework, and locate the taking of the joint action in the road traffic state as the output of the Critic network;
[0090] Construct a joint action space according to the defined Actor-Critic framework.
[0091] The construction process of the joint action space is based on the deep fusion and collaborative optimization of multi-dimensional traffic elements. The electronic device can first generate a discrete signal control agent with the green light duration and phase switching transition time in the traffic signal as the core parameters. This agent can accurately describe the timing characteristics of signal timing and its impact on traffic flow. In the dimension of vehicle motion control, the electronic device uses the steering angle and acceleration as continuous control variables to generate a continuous path planning agent to achieve refined modeling of the vehicle driving trajectory. Through the Cartesian product operation, the discrete signal control agent and the continuous path planning agent are mathematically mapped to form a unified hybrid action space, thus completely describing the collaborative relationship between traffic signals and vehicle motion. Among them, the hybrid action space is specifically the joint action of the space.
[0092] In the specific implementation of the Actor-Critic framework, the electronic device defines the joint action in the hybrid action space as the output of the Actor network. This Actor network simultaneously generates a signal control strategy and a path planning strategy through policy optimization. At the same time, the electronic device takes the road traffic state and the joint action as the input of the Critic network. Among them, the road traffic state includes traffic flow density, vehicle speed distribution, queue length, etc., and evaluates the global utility of the current strategy through the value function. The output of the Critic network reflects the long-term benefit of taking the joint action in this road traffic state, provides the gradient direction for the policy update of the Actor network, and the electronic device can locate the taking of the joint action in the road traffic state as the output of this Critic network. The electronic device constructs a joint action space according to the defined Actor-Critic framework.
[0093] In one example, the electronic device can define the signal control agent as a discrete action space A signal∈ {P1, P2, ……, P n}, where each phase P i includes a green duration (5 - 60 seconds) and a switching transition time (3 - 5 seconds). The path planning agent can be defined as a continuous action space A vehicle ∈ R 2 , where R includes a steering angle θ ∈ [-30°, +30°] and an acceleration a ∈ [-3m / s, +3m / s]. The hybrid action space can be expressed as A joint = A signal × A vehicle , and the dimension is expanded to n + 2.
[0094] To obtain the joint action space, based on the above embodiments, in the embodiments of the present invention, the pre-training process of the joint action space includes:
[0095] Regarding the road traffic state in the joint action space to be trained as the state quantity of the autonomous driving path planning scenario;
[0096] Regarding the joint action in the hybrid action space as the action quantity of the autonomous driving path planning scenario;
[0097] Based on the state quantity and the action quantity, using the improved MADDPG algorithm to perform collaborative iterative training on the joint action space to be trained, and obtaining the trained joint action space.
[0098] In the embodiments of the present invention, the pre-training process of the joint action space is implemented by constructing a simulation environment for autonomous driving path planning. The electronic device can first regard the road traffic state in the joint action space to be trained as the state quantity in the autonomous driving planning scenario. This road traffic state includes information such as traffic flow density, signal phase, vehicle position, etc. This state quantity is used to comprehensively describe the dynamic characteristics of the traffic scenario. And it can define the joint action in the hybrid action space as the action quantity of the autonomous driving path planning scenario. This joint action is a combination of a signal control strategy and a vehicle path planning strategy. This action quantity can be used as the decision-making output for the interaction between the agent and the environment.
[0099] Based on this state quantity and this action quantity, using the improved MADDPG algorithm to perform coordinated iterative training on the joint action space to be trained, and obtaining the trained joint action space. Among them, in the training stage, the Critic network uses the global state and action information to evaluate the value of the joint strategy and guides the strategy update of the Actor network; in the execution stage, each agent makes independent decisions only relying on local observation information. Through multiple rounds of iterative training, the electronic device gradually optimizes the collaborative strategy of signal control and path planning, and finally obtains a joint action space that can efficiently handle complex traffic scenarios. This pre-training mechanism significantly improves the convergence speed and policy stability.
[0100] To obtain the trained joint action space, based on the above embodiments, in the embodiments of the present invention, based on the state quantity and the action quantity, the improved MADDPG algorithm is used to perform collaborative iterative training on the joint action space to be trained, and the obtained trained joint action space includes:
[0101] Based on the improved MADDPG algorithm, under the constraint of the spatio-temporal coupling matrix, the parameters in the joint action space to be trained are initialized; the parameters include the network parameters of the Actor-Critic framework.
[0102] Based on the state quantity, the Actor-Critic framework with initialized parameters is used to generate the corresponding action quantity, and the value estimation value when the signal control agent and the path planning agent jointly execute the action quantity is calculated.
[0103] Based on the value estimation value, the network parameters of the Critic network and the Actor network in the Actor-Critic framework are iteratively optimized to obtain the trained joint action space.
[0104] In the embodiments of the present invention, the electronic device can initialize the parameters in the joint action space to be trained under the constraint of the spatio-temporal coupling matrix. The parameters in the joint action space to be trained include the weights, biases and other parameters of the Critic network and the Actor network in the Actor-Critic framework.
[0105] The electronic device uses the initialized Actor network to generate the joint action quantity according to the current traffic state, including the signal control strategy and the path planning strategy, and evaluates the global value estimation value of the joint action through the Critic network. The value estimation value comprehensively considers multiple indicators such as traffic efficiency, signal coordination and vehicle driving cost.
[0106] In the optimization stage, the electronic device can synchronously update the parameters of the Critic network and the Actor network based on the value estimation value through the gradient backpropagation algorithm or other algorithms: the Critic network improves the accuracy of value estimation by minimizing the temporal difference error, and the Actor network optimizes the generation strategy of the joint action through the policy gradient method. Through multiple rounds of iterative training, it gradually converges to the optimal strategy, and finally obtains a joint action space that can efficiently coordinate signal control and path planning, providing a reliable solution for collaborative optimization in complex traffic scenarios.
[0107] In one example, the electronic device can copy the weights of the main network to the target network regularly or at each step with a small learning rate to maintain the stability of training.
[0108] In order to obtain a trained joint action space, based on the above embodiments, in the embodiments of the present invention, the network parameters of the Critic network and the Actor network in the Actor-Critic framework are iteratively optimized based on the value estimate, and the obtained trained joint action space includes:
[0109] Based on the value estimate, calculate the loss value using the mean squared error loss function; optimize the network parameters of the Critic network based on the loss value;
[0110] Use road simulation technology to simulate the combined action amounts executed by the signal control agent and the path planning agent to obtain simulated traffic data;
[0111] Based on the simulated traffic data, calculate the reward values of the signal control agent and the path planning agent using the reward function; optimize the network parameters of the Actor network based on the reward values;
[0112] Based on the Critic network and the Actor network with optimized network parameters, continue to generate the corresponding action amounts until the iteration termination condition is met, and obtain the trained joint action space.
[0113] In the embodiments of the present invention, the electronic device can estimate the value of the current policy through the Critic network and use the mean squared error loss function to calculate the difference between the predicted value and the actual value, and use this as the loss value. Based on this loss value, update the parameters of the Critic network. In one example, optimization algorithms such as gradient descent can be used to update the parameters of the Critic network to improve the accuracy of its value estimation.
[0114] With the help of road simulation technology, simulate the combined action amounts executed by the signal control agent and the path planning agent to obtain simulated traffic data. This process can generate rich simulated traffic data, including key information such as vehicle flow, driving speed, and signal light status.
[0115] Based on the obtained simulated traffic data, the electronic device can use the reward function to quantitatively evaluate the performance of the signal control agent and the path planning agent. The reward function comprehensively considers multiple dimensions such as traffic efficiency, safety, and environmental protection to ensure that the behavior of the agent meets the requirements of actual traffic management. According to the calculated reward values, further optimize the parameters of the Actor network.
[0116] After obtaining the Critic network and the Actor network with optimized network parameters, the optimized Critic network and Actor network continue to generate the corresponding action amounts until the iteration termination condition is met, and the trained joint action space can be obtained.
[0117] To obtain the trained joint action space, based on the above embodiments, in the embodiments of the present invention, the reward function satisfies the following formula:
[0118]
[0119] Wherein, R total is the reward value, α is a preset weight, X is the overall traffic efficiency of a specific area, N is the number of vehicles in the specific area, and x i is the actual driving time of the i-th vehicle in the driving data, and is the expected driving time of the i-th vehicle in the driving data.
[0120] The reward function provided by the embodiments of the present invention comprehensively considers factors such as global traffic efficiency and individual vehicle travel time deviation. This reward structure encourages the signal control agent and the path planning agent to not only pursue local optimal solutions but also consider the overall performance.
[0121] To accurately and effectively perform autonomous driving planning, based on the above embodiments, in the embodiments of the present invention, before controlling the traffic lights and vehicles in a specific area using the optimal signal phase switching time and vehicle driving path, it further includes:
[0122] Based on the optimal signal phase switching time and vehicle driving path in the optimal joint action, use the conflict detection function to determine the traffic conflict detection result;
[0123] If the traffic conflict detection result indicates the existence of potential traffic conflicts, then use a preset correction mechanism to correct the optimal joint action.
[0124] Before implementing the optimal signal phase switching time and vehicle driving path to regulate the traffic lights and vehicles in a specific area, the electronic device also needs to ensure smooth and safe traffic. Specifically, the electronic device can use the conflict detection function to identify potential traffic conflicts based on the determined optimal joint action and determine the traffic conflict detection result.
[0125] If the traffic conflict detection result indicates the existence of potential traffic conflicts, then the electronic device can use a preset correction mechanism to correct the optimal joint action. This mechanism can quickly adjust the signal phase switching time and vehicle driving path to avoid conflicts, thus ensuring the harmony and efficiency of the traffic system.
[0126] In the embodiments of the present invention, the electronic device uses the conflict detection function to identify possible conflicts and triggers the corresponding correction mechanism when a conflict is detected, ensuring traffic safety and efficiency.
[0127] Suppose at an intersection, the traffic signal has planned specific phase switching times according to an optimal joint action plan, and at the same time, vehicles have selected optimal driving paths based on the suggestions of path planning agents. However, the conflict detection function finds that there may be an intersection conflict between the straight-going vehicles in the east-west direction (where east-west refers to the actual east-west in the scenario) and the left-turning vehicles in the north-south direction (where north-south refers to the actual north-south in the scenario) (where left and right refer to the actual left and right in the scenario) at a certain point in time. Then the preset correction mechanism can be to extend the green light time in the east-west direction to allow the straight-going vehicles to pass through the intersection as soon as possible. Delay the start time of the left-turn green light in the north-south direction to avoid entering the intersection at the same time as the vehicles in the east-west direction.
[0128] In the embodiment of the present invention, the spatio-temporal coupling matrix provides necessary constraint conditions, enabling the signal control agent and the path planning agent to fully consider the time dynamic characteristics of traffic flow when making decisions; while the joint action space allows the signal control agent and the path planning agent to explore more complex cooperation modes. Combining with the improved MADDPG algorithm, it can effectively promote the collaborative optimization of the signal control agent and the path planning agent in complex traffic scenarios.
[0129] The present invention discloses a traffic flow optimization and autonomous driving path planning method based on reinforcement learning. By constructing an intelligent simulation system of the traffic environment, the traffic signal timing and the path planning of autonomous driving vehicles are dynamically adjusted using reinforcement learning algorithms. It includes a traffic data collection module, a dynamic environment modeling module, an intelligent decision-making module, and a multi-objective optimization module. Through the multi-agent reinforcement learning algorithm (Multi-Agent Reinforcement Learning, MARL), it can achieve global optimization of traffic flow in complex traffic scenarios, reduce congestion, and improve the traffic efficiency of autonomous driving vehicles. The present invention is particularly suitable for the development of intelligent traffic management systems and autonomous driving technologies, providing an innovative solution for future traffic systems.
[0130] Embodiment 2:
[0131] Based on the same inventive concept, the present invention also provides an autonomous driving path planning system, the structural schematic diagram is as Figure 2 shown, including:
[0132] A traffic data collection module 201, configured to obtain road traffic data of a specific area;
[0133] A dynamic environment modeling module 202, configured to solve the joint action space of the signal control agent and the path planning agent based on the road traffic data in combination with the spatio-temporal coupling matrix to obtain an optimal joint action; the optimal joint action includes the optimal signal phase switching time and vehicle driving path in the specific area;
[0134] The intelligent decision-making module 203 is used to control the traffic lights and vehicles in the specific area by using the optimal signal phase switching time and the vehicle driving path; the joint action space is constructed by forming a hybrid action space through the Cartesian product of the signal control agent and the path planning agent, and then combining the hybrid action space with the Actor-Critic framework.
[0135] In a specific implementation manner, the dynamic environment modeling module 202 is further used to construct the joint action space in the following way:
[0136] Generate discrete signal control agents according to the green light duration and switching transition time in the traffic lights;
[0137] Generate continuous path planning agents according to the steering angle and acceleration during vehicle driving;
[0138] Form a hybrid action space through the Cartesian product of the signal control agent and the path planning agent;
[0139] Define the joint action in the hybrid action space as the output of the Actor network in the Actor-Critic framework;
[0140] Define the road traffic state and the joint action as the input of the Critic network in the Actor-Critic framework, and locate the joint action taken under the road traffic state as the output of the Critic network;
[0141] Construct the joint action space according to the defined Actor-Critic framework.
[0142] In a specific implementation manner, it further includes a multi-objective optimization module 204, which is used to train the joint action space in the following way:
[0143] Take the road traffic state in the joint action space to be trained as the state quantity of the autonomous driving path planning scenario;
[0144] Take the joint action in the hybrid action space as the action quantity of the autonomous driving path planning scenario;
[0145] Based on the state quantity and the action quantity, use the improved multi-agent deep deterministic policy gradient MADDPG algorithm to perform collaborative iterative training on the joint action space to be trained, and obtain the trained joint action space.
[0146] In a specific implementation manner, the multi-objective optimization module 204 is specifically configured to initialize the parameters in the joint action space to be trained based on an improved MADDPG algorithm under the constraint of a spatio-temporal coupling matrix; the parameters include the network parameters of the Actor-Critic framework.
[0147] Based on the state quantity, use the Actor-Critic framework with initialized parameters to generate corresponding action quantities, and calculate the value estimation value when the signal control agent and the path planning agent jointly execute the action quantities.
[0148] Based on the value estimation value, iteratively optimize the network parameters of the Critic network and the Actor network in the Actor-Critic framework to obtain a trained joint action space.
[0149] In a specific implementation manner, the multi-objective optimization module 204 is specifically configured to calculate a loss value based on the value estimation value using a mean squared error loss function; optimize the network parameters of the Critic network based on the loss value.
[0150] Use road simulation technology to simulate the joint execution of the action quantities by the signal control agent and the path planning agent to obtain simulated traffic data.
[0151] Based on the simulated traffic data, calculate the reward values of the signal control agent and the path planning agent using a reward function; optimize the network parameters of the Actor network based on the reward values.
[0152] Based on the Critic network and the Actor network with optimized network parameters, continue to generate corresponding action quantities until the iteration termination condition is met to obtain a trained joint action space.
[0153] In a specific implementation manner, the intelligent decision-making module 203 is further configured to determine a traffic conflict detection result using a conflict detection function based on the best signal phase switching time and vehicle driving path in the best joint action.
[0154] If the traffic conflict detection result indicates the existence of potential traffic conflicts, a preset correction mechanism is used to correct the best joint action.
[0155] Embodiment 3:
[0156] As Figure 3As shown, the present invention also provides an electronic device, which may be a computer device, a single-chip microcomputer device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, the processor, and the transceiver component are connected through a bus; the memory can be used to store an execution program, and an exemplary execution program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, and this data can be called and / or modified when the instructions are executed.
[0157] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of an automatic driving path planning method in the above embodiment.
[0158] Embodiment 4:
[0159] Based on the same inventive concept, the present invention also provides a readable storage medium, specifically an electronic device-readable storage medium (Memory). The electronic device-readable storage medium is a memory device in the electronic device and is used to store programs and data. It can be understood that the storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more execution programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. By the processor loading and executing one or more instructions stored in the storage medium, the steps of an automatic driving path planning method in the above embodiment can be implemented.
[0160] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0161] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more of the processes or Figure 1 blocks or the combination of blocks.
[0162] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in Figure 1 one or more of the processes or Figure 1 blocks or the combination of blocks.
[0163] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one or more of the processes or Figure 1 blocks or the combination of blocks.
[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the scope of its protection. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that after reading the present invention, various changes, modifications, or equivalent replacements can still be made to the specific implementation manners of the application. However, these changes, modifications, or equivalent replacements are all within the scope of the protection of the pending claims of the application.
Claims
1. An autonomous driving path planning method, characterized in that, The method includes: Obtaining the road traffic data of a specific area; Based on the road traffic data and combined with a spatio-temporal coupling matrix, solving the joint action space of a signal control agent and a path planning agent to obtain an optimal joint action; the optimal joint action includes the optimal signal phase switching time and vehicle driving path within the specific area; Using the optimal signal phase switching time and vehicle driving path to control the traffic lights and vehicles within the specific area; The joint action space is formed by taking the Cartesian product of the signal control agent and the path planning agent to form a hybrid action space, and then constructing the hybrid action space in combination with the Actor-Critic framework.
2. The method according to claim 1, wherein The construction process of the joint action space includes: Generating a discrete signal control agent according to the green light duration and switching transition time in the traffic lights; Generating a continuous path planning agent according to the steering angle and acceleration during vehicle driving; Forming a hybrid action space by taking the Cartesian product of the signal control agent and the path planning agent; Defining the joint action in the hybrid action space as the output of the Actor network in the Actor-Critic framework; Defining the road traffic state and the joint action as the input of the Critic network in the Actor-Critic framework, and positioning the taking of the joint action under the road traffic state as the output of the Critic network; Constructing a joint action space according to the defined Actor-Critic framework.
3. The method according to claim 1 or 2, characterized in that, The pre-training process of the joint action space includes: Regarding the road traffic state in the joint action space to be trained as the state quantity of the autonomous driving path planning scenario; Regarding the joint action in the hybrid action space as the action quantity of the autonomous driving path planning scenario; Based on the state quantity and the action quantity, using an improved Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm to perform collaborative iterative training on the joint action space to be trained, and obtaining a trained joint action space.
4. The method according to claim 3, wherein The performing collaborative iterative training on the joint action space to be trained based on the state quantity and the action quantity using the improved MADDPG algorithm to obtain a trained joint action space includes: Based on the improved MADDPG algorithm, initializing the parameters in the joint action space to be trained under the constraint of the spatio-temporal coupling matrix; the parameters include the network parameters of the Actor-Critic framework; Based on the state quantity, using the Actor-Critic framework with initialized parameters to generate corresponding action quantities, and calculating the value estimation value when the signal control agent and the path planning agent jointly execute the action quantities; Iteratively optimizing the network parameters of the Critic network and the Actor network in the Actor-Critic framework based on the value estimation value to obtain a trained joint action space.
5. The method according to claim 4, wherein Iteratively optimizing the network parameters of the Critic network and the Actor network in the Actor-Critic framework based on the value estimate to obtain a trained joint action space includes: Based on the value estimate, calculating a loss value using a mean squared error loss function; optimizing the network parameters of the Critic network based on the loss value; Using road simulation technology to simulate the joint execution of the action amounts by the signal control agent and the path planning agent to obtain simulated traffic data; Based on the simulated traffic data, calculating the reward values of the signal control agent and the path planning agent using a reward function; optimizing the network parameters of the Actor network based on the reward values; Continuing to generate corresponding action amounts based on the Critic network and the Actor network with optimized network parameters until an iteration termination condition is met, to obtain a trained joint action space.
6. The method according to claim 5, characterized in that, The reward function satisfies the following formula: Among them, R total is the reward value, α is a preset weight, X is the overall traffic efficiency of the specific area, N is the number of vehicles in the specific area, and x i is the actual travel time of the i-th vehicle in the travel data, and is the expected travel time of the i-th vehicle in the travel data.
7. The method according to claim 1, wherein Before using the optimal signal phase switching time and vehicle driving path to control the traffic lights and vehicles in the specific area, it further includes: Based on the optimal signal phase switching time and vehicle driving path in the optimal joint action, using a conflict detection function to determine a traffic conflict detection result; If the traffic conflict detection result indicates the existence of potential traffic conflicts, a preset correction mechanism is used to correct the optimal joint action.
8. An automatic driving path planning system, characterized in that, It includes: A traffic data collection module for acquiring road traffic data of a specific area; A dynamic environment modeling module for solving the joint action space of the signal control agent and the path planning agent based on the road traffic data combined with a spatio-temporal coupling matrix to obtain an optimal joint action; the optimal joint action includes the optimal signal phase switching time and vehicle driving path in the specific area; An intelligent decision-making module for using the optimal signal phase switching time and vehicle driving path to control the traffic lights and vehicles in the specific area; the joint action space is formed by taking the Cartesian product of the signal control agent and the path planning agent to form a hybrid action space, and then constructing it in combination with the Actor-Critic framework.
9. An electronic device, characterized in that, It includes: At least one processor and a memory; The memory and the processor are connected by a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the autonomous driving path planning method according to any one of claims 1-7 is implemented.
10. A readable storage medium, characterized in that, There is an execution program stored thereon, and when the execution program is executed, the autonomous driving path planning method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Reinforcement learning area signal control method based on vehicle planning path
CN113487902A
Method and system for path navigation in dynamic environment based on deep reinforcement learning
CN116242379A
Non-signalized intersection vehicle passing decision planning method, system and equipment
CN118212808A
Method and device for the computer-implemented routing of motor vehicles in a predetermined area
EP3723062A1
Individual interactive energy-optimized routing of traffic participants
WO2012059275A1