Vehicle path planning method, terminal equipment and storage medium
By using the Monte Carlo search tree method in path planning, the state search tree of vehicles and obstacle vehicles is constructed, and the UCT value is calculated to determine the decision branch, the path difference problem caused by the uncertainty of obstacle vehicles in the prior art is solved, and driving safety is improved.
Patent Information
- Application Number
- CN202411855288.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-06
AI Technical Summary
When the existing path planning method deals with the uncertainty of the movement of the obstacle vehicle, the predicted path and the actual path are different, resulting in the possibility of collision between the obstacle vehicle.
Using the Monte Carlo search tree method, by constructing a search tree for the current state of the vehicle and the obstacle vehicle state, the UCT value of each node is calculated to determine the decision branch, and the best decision branch is selected to control the vehicle's driving.
Through the Monte Carlo search tree method, the future status of the vehicle can be simulated more accurately, reducing the risk of collision with obstacle vehicles, and improving driving safety.
Smart Images

Figure CN119935135A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent driving technology, and in particular to a vehicle path planning method, terminal equipment, and storage medium. Background Art
[0002] In path planning related technologies, deep learning is often used to predict the running paths of obstacle vehicles around the current vehicle. Based on these running paths, the current vehicle plans a driving path that does not collide with the obstacle vehicles. However, due to the uncertainty of the action of the obstacle vehicle, the running path predicted based on the current action of the obstacle vehicle is significantly different from the actual running path, resulting in the current vehicle running along the planned driving path, and there is a possibility of collision with the obstacle vehicle. Summary of the invention
[0003] The present application provides a vehicle path planning method, terminal device and storage medium.
[0004] A technical solution adopted in the present application is to provide a vehicle path planning method, the method comprising:
[0005] During the driving process of the vehicle, obtain the obstacle vehicle corresponding to the current vehicle;
[0006] According to the first state of the current vehicle and the second state of the obstacle vehicle, a Monte Carlo search tree is constructed; the first state includes the current position, current speed and current vehicle head direction of the current vehicle at the current moment; the second state includes the current position, current speed and current vehicle head direction of the obstacle vehicle at the current moment; wherein the first state is the root node of the Monte Carlo search tree;
[0007] According to the UCT value corresponding to each node in the Monte Carlo search tree, several decision branches are determined, and the best decision branch is determined from the several decision branches. The node is used to simulate the running state of the current vehicle at a future moment;
[0008] Control the current vehicle driving according to the best decision branch.
[0009] Optionally, a Monte Carlo search tree is constructed according to the first state of the current vehicle and the second state of the obstacle vehicle, including:
[0010] The root node is used as the parent node to generate at least one child node; the child node is used to simulate the first target state of the current vehicle after the first preset time; and, according to the second state, update the second target state of the obstacle vehicle after the first preset time; wherein the current vehicle in the first target state does not collide with the obstacle vehicle in the second target state;
[0011] Taking the child node as the parent node, and continuing to execute the step of generating at least one child node until the number of layers of the Monte Carlo search tree is greater than or equal to the preset number of layers;
[0012] When the number of layers of the Monte Carlo search tree is greater than or equal to the preset number of layers, the score corresponding to each node is calculated, and the UCT value corresponding to each node is determined using the score.
[0013] Optionally, the first state includes a current acceleration of the current vehicle;
[0014] The first target state includes a target acceleration, a first target heading angle, and a first target curvature of the current vehicle, and the score includes a driving comfort level;
[0015] Calculate the score corresponding to each node, including:
[0016] Using the current acceleration and the target acceleration, the acceleration comfort corresponding to the current node is obtained;
[0017] Using the first target heading angle and the second target heading angle of the parent node corresponding to the current node, the heading angle comfort corresponding to the current node is obtained;
[0018] Using the first target curvature and the second target curvature of the parent node, the curvature comfort corresponding to the current node is obtained;
[0019] The weighted sum of acceleration comfort, heading angle comfort, and curvature comfort is calculated to obtain the driving comfort of the current node.
[0020] Optionally, the first target state includes a target speed and a first target position;
[0021] The scores include traffic efficiency;
[0022] Calculate the score corresponding to each node, including:
[0023] Using the target speed and the maximum speed limit of the road where the vehicle is currently located, the speed efficiency corresponding to the current node is obtained;
[0024] Get all associated nodes between the root node and the current node, and use the corresponding position information of all associated nodes to determine the target path length;
[0025] The target longitudinal distance and the target path length are used to obtain the traffic efficiency of the position corresponding to the current node; the target longitudinal distance is the longitudinal distance between the first target position and the second target position of the parent node corresponding to the current node;
[0026] Calculate the weighted sum of speed efficiency and position efficiency to get the efficiency of the current node.
[0027] Optionally, the first target state includes a target acceleration;
[0028] The scores include historical scores;
[0029] Calculate the score corresponding to each node, including:
[0030] Obtain the historical acceleration of the current vehicle at the historical moment, and determine the preset acceleration range based on the historical acceleration;
[0031] In response to the target acceleration being within the preset acceleration interval, a first preset value is added to the history score.
[0032] Optionally, the first target state includes a first target position;
[0033] The score includes a lane-changing penalty;
[0034] Calculate the score corresponding to each node, including:
[0035] According to the second target position of the target node corresponding to the parent node and the first target position, determining whether the current vehicle changes lanes;
[0036] In response to the current vehicle having a lane change behavior, a second preset value is added to the lane change penalty item.
[0037] Optionally, several decision branches are determined according to the UCT value corresponding to each node in the Monte Carlo search tree, including:
[0038] Use the root node to build a decision branch;
[0039] Take the root node as the parent node, and add the target child node with the largest UCT value from the child nodes of the next layer to the decision branch;
[0040] Take the target child node as the parent node, and repeatedly add the target child node with the largest UCT value from the next layer of child nodes to the decision branch until the number of layers where the target child node is located is greater than or equal to the preset number of layers.
[0041] Optionally, during the driving process of the vehicle, obtaining an obstacle vehicle corresponding to the current vehicle includes:
[0042] Get candidate vehicles around the current vehicle;
[0043] In response to a collision between the candidate vehicle and the current vehicle within a second preset time period, the candidate vehicle is regarded as an obstacle vehicle.
[0044] Another technical solution adopted by the present application is to provide a terminal device, the terminal device comprising a memory and a processor connected to the memory;
[0045] The memory is used to store program data, and the processor is used to execute the program data to implement the vehicle path planning method as described above.
[0046] Another technical solution adopted by the present application is to provide a computer storage medium, wherein the computer storage medium is used to store program data, and when the program data is executed by a computer, it is used to implement the vehicle path planning method as described above.
[0047] The beneficial effects of the present application are: during vehicle driving, an obstacle vehicle corresponding to the current vehicle is obtained, and a Monte Carlo search tree is constructed based on the first state of the current vehicle and the second state of the obstacle vehicle; a number of decision branches are determined based on the UCT value corresponding to each node in the Monte Carlo search tree, and then the current vehicle is controlled to drive based on the best decision branch. The path planning method provided by the present application uses the state of the current vehicle at the current moment as the root node, performs a Monte Carlo tree search on the future running state of the current vehicle, and adds the driving state with the highest UCT value to the best decision branch. The best decision branch obtained can reduce the possibility of collision between the current vehicle and the obstacle vehicle, and improve the driving safety of the current vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 It is a flow chart of an embodiment of a vehicle path planning method provided by the present application;
[0050] Figure 2 is a flow chart of another embodiment of the vehicle path planning method provided by the present application;
[0051] Figure 3 is a schematic diagram of an embodiment of a self-vehicle and an interactive vehicle;
[0052] Figure 4 It is a structural diagram of an embodiment of a terminal device provided by the present application;
[0053] Figure 5 It is a structural diagram of an embodiment of a computer storage medium provided by the present application. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0055] It is not easy for smart driving cars to deal with the problem of interacting with other obstacle vehicles around them. Especially in the scenario where the action intention of the surrounding obstacles is uncertain, or the scenario where the multimodal predicted action probabilities of the surrounding obstacles obtained through data drive are similar, it is often unreasonable for the smart driving car to take the action with the highest probability as the final prediction information of the obstacle vehicle. It is possible that in the next cycle, the predicted probability of the obstacle vehicle for the action will change significantly; the action with a slightly lower probability in the previous cycle will have a significantly higher probability than other actions at this time, so the trajectory planned by the smart vehicle based on the action predicted in the previous cycle is usually unreasonable, and even a large jump will occur.
[0056] Therefore, we need to comprehensively consider all possible actions of obstacles at the very beginning, and then make corresponding plans. The method proposed in this paper is based on the Monte Carlo search tree. The first part is that the actions searched in the common action interval are consistent. For all possible actions of obstacles, the planning made by the vehicle can deal with them well. In the next cycle or after the obstacle has a clear intention, the second part is to select a specific action branch in the specific situation action interval according to the predicted action, so as to better match the action and make the trajectory planning more reasonable.
[0057] This application mainly designs a set of methods to improve the rationality of vehicle path planning. Unlike traditional methods, this application starts from the perspective of constructing a Monte Carlo search tree, deduces the different motion states of the current vehicle at future moments, and then selects the best decision branch from the constructed Monte Carlo search tree, and controls the current vehicle driving according to the best decision branch to generate a more reasonable path trajectory.
[0058] Please refer to Figure 1 , Figure 1 It is a flow chart of an embodiment of a vehicle path planning method provided in the present application.
[0059] like Figure 1 As shown, the vehicle path planning method of the embodiment of the present application may specifically include the following steps:
[0060] S1, during the driving process of the vehicle, obtaining an obstacle vehicle corresponding to the current vehicle.
[0061] The path planning method for the vehicle provided in the present application is mainly performed by a path planning device. In some embodiments, the path planning device may be the domain controller of the current vehicle itself. In some embodiments, the turning angle determination device may be a device that is communicatively connected to the domain controller. For example, the device may be a device for monitoring images, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, and an intelligent driving car, a robot, a security system, and any one or more products of glasses and helmets for augmented reality or virtual reality. In some possible implementations, the path planning method for the vehicle may be implemented by a processor calling computer-readable instructions stored in a memory.
[0062] Specifically, during the driving process of the vehicle, the path planning device uses the on-board sensors on the current vehicle to obtain the obstacle vehicle corresponding to the current vehicle.
[0063] In some embodiments, the vehicle-mounted sensor may include at least one of a vehicle-mounted camera, a lidar, a millimeter-wave radar, and an ultrasonic radar.
[0064] In some embodiments, the current vehicle is an intelligent driving vehicle.
[0065] In some embodiments, the obstacle vehicle refers to a vehicle that collides with the current vehicle.
[0066] Among some possible application scenarios, this embodiment is mainly aimed at the scenario where the current vehicle changes lanes or merges and splits, and the obstacle vehicle can be a vehicle located in front of or behind the current vehicle, or a vehicle located on the left or right of the current vehicle.
[0067] S2: constructing a Monte Carlo search tree according to the first state of the current vehicle and the second state of the obstacle vehicle.
[0068] The first state includes the current position, current speed and current front direction of the current vehicle at the current moment; the second state includes the current position, current speed and current front direction of the obstacle vehicle at the current moment; wherein the first state is the root node of the Monte Carlo search tree.
[0069] Specifically, the path planning device constructs the root node of the Monte Carlo search tree using the current position, current speed, and current vehicle head direction at the current moment.
[0070] Optionally, the current position may be represented by a horizontal coordinate and a vertical coordinate of a global coordinate system, or may be represented by a horizontal coordinate and a vertical coordinate of a Frenet coordinate system.
[0071] In some possible embodiments, the path planning state utilizes the current position, current speed, current acceleration, current vehicle head direction, and size information of the current vehicle at the current moment to construct the root node of the Monte Carlo search tree.
[0072] In some possible embodiments, the second state may include the current position, current speed, current acceleration, current vehicle head direction and size information of the obstacle vehicle at the current moment.
[0073] S3, determining a plurality of decision branches according to the UCT value corresponding to each node in the Monte Carlo search tree, and determining the best decision branch from the plurality of decision branches.
[0074] The node is used to simulate the operating state of the current vehicle at a future moment.
[0075] In some embodiments, the node can be used to simulate the operating state of the current vehicle at a future moment without colliding with any obstacle vehicle.
[0076] In some embodiments, the path search device determines a plurality of decision branches based on the size of the UCT value corresponding to each node in the Monte Carlo search tree, and determines the best decision branch from the plurality of decision branches. Exemplarily, the path search device calculates the sum of the UCT values corresponding to each decision branch, and takes the decision branch with the largest sum of UCT values as the best decision branch.
[0077] In some possible embodiments, the operating state of the current vehicle at a future moment includes the position, speed, acceleration, and vehicle head direction of the current vehicle at a future moment.
[0078] S4, controlling the current vehicle to travel according to the optimal decision branch.
[0079] Specifically, the path planning device controls the current vehicle driving according to the best decision branch.
[0080] It can be understood that each node in the optimal decision branch corresponds to the operating state of the current vehicle at a future moment, and the path search device can generate a control instruction based on the above operating state and the first state to control the driving of the current vehicle.
[0081] In the above scheme, during the driving process of the vehicle, the obstacle vehicle corresponding to the current vehicle is obtained, and a Monte Carlo search tree is constructed based on the first state of the current vehicle and the second state of the obstacle vehicle; a number of decision branches are determined according to the UCT value corresponding to each node in the Monte Carlo search tree, and then the driving of the current vehicle is controlled according to the best decision branch. The path planning method provided by the present application uses the state of the current vehicle at the current moment as the root node, performs a Monte Carlo tree search on the future running state of the current vehicle, and adds the driving state with the highest UCT value to the best decision branch. The best decision branch obtained can reduce the possibility of collision between the current vehicle and the obstacle vehicle and improve the driving safety of the current vehicle.
[0082] Another embodiment of the vehicle path planning method provided by the present application may specifically include the following steps:
[0083] S11, during the driving process of the vehicle, obtaining an obstacle vehicle corresponding to the current vehicle.
[0084] S12, taking the root node as the parent node and generating at least one child node.
[0085] The subnode is used to simulate the first target state of the current vehicle after the first preset time period.
[0086] Specifically, the path planning device takes the root node as a parent node and generates at least one child node.
[0087] In this step, these sub-nodes are used to simulate the running state of the current vehicle after a first preset time period from the current moment.
[0088] Exemplarily, the first preset time length may be any one of 1s, 0.5s, and 0.1s.
[0089] In some embodiments, the path planning device also updates the second target state of the obstructing vehicle after the first preset time period according to the second state corresponding to the obstructing vehicle. For example, the path planning device updates the position information and speed information of the obstructing vehicle after the first preset time period according to the current position, current speed, current acceleration, and current vehicle head direction corresponding to the obstructing vehicle.
[0090] Among them, the current vehicle in the first target state does not collide with the obstacle vehicle in the second target state.
[0091] In some embodiments, the path planning device generates a first maximum circumscribed cuboid representing the current vehicle based on the maximum length, maximum width and maximum height of the current vehicle. Similarly, the path planning device generates a second maximum circumscribed cuboid representing the obstacle vehicle based on the maximum length, maximum width and maximum height of the obstacle vehicle.
[0092] Furthermore, the path planning device determines whether the first largest circumscribed cuboid in the first target state contacts the second largest circumscribed cuboid in the second target state. If so, it means that the current vehicle in the first target state collides with the obstacle vehicle in the second target state.
[0093] It can be understood that the layer where the child node in this step is located is the first layer.
[0094] S13, taking the child node as the parent node, and continuing to execute the step of generating at least one child node until the number of layers of the Monte Carlo search tree is greater than or equal to the preset number of layers.
[0095] Specifically, the path planning device uses the child node of the first layer as a parent node to generate at least one child node in the second layer. Then, the child node of the second layer is used as a parent node to generate at least one child node in the third layer. The child nodes are generated repeatedly until the number of layers of the Monte Carlo search tree is greater than or equal to the preset number of layers.
[0096] It is understandable that the more node layers the Monte Carlo search tree has, the greater the computational complexity. In addition, due to the uncertainty of the road environment, if the number of node layers of the Monte Carlo search tree is too large, the current vehicle driving according to the best decision branch may not be consistent with the future road conditions. Therefore, this application has certain restrictions on the number of node layers of the Monte Carlo search tree.
[0097] Exemplarily, the preset number of layers may be any one of 5 layers, 6 layers, 7 layers, 8 layers, and 9 layers, which is not limited here.
[0098] In some possible embodiments, a node of a certain layer is used to represent the first target state of the current vehicle within a certain 1s.
[0099] S14, when the number of layers of the Monte Carlo search tree is greater than or equal to the preset number of layers, the score corresponding to each node is calculated, and the UCT value corresponding to each node is determined using the score.
[0100] In some possible embodiments, the score may be determined based on the first target state of the current vehicle.
[0101] In some possible embodiments, the score may be determined based on the first target state and the second target state.
[0102] In some embodiments, the UCT value of a certain node m satisfies the following relationship:
[0103]
[0104] Where Q(m) is the average score of node m, N is the number of visits to the parent node of node m, n is the number of visits to node m, and c is a constant that controls the balance between exploration and exploitation.
[0105] In some embodiments, Q(m) is the quotient of the score of node m and the number of visits to node m, that is, it satisfies the following relationship:
[0106]
[0107] Among them, cost_m is the score of node m.
[0108] It should be noted that, in the embodiment of the present application, the higher the score, the more the status of the node meets the requirements.
[0109] S15, determining a plurality of decision branches according to the UCT value corresponding to each node in the Monte Carlo search tree, and determining the best decision branch from the plurality of decision branches.
[0110] S16, controlling the current vehicle driving according to the best decision branch.
[0111] In the above scheme, during the driving process of the vehicle, the obstacle vehicle corresponding to the current vehicle is obtained, and a Monte Carlo search tree is constructed based on the first state of the current vehicle and the second state of the obstacle vehicle; a number of decision branches are determined according to the UCT value corresponding to each node in the Monte Carlo search tree, and then the driving of the current vehicle is controlled according to the best decision branch. The path planning method provided by the present application uses the state of the current vehicle at the current moment as the root node, performs a Monte Carlo tree search on the future running state of the current vehicle, and adds the driving state with the highest UCT value to the best decision branch. The best decision branch obtained can reduce the possibility of collision between the current vehicle and the obstacle vehicle and improve the driving safety of the current vehicle.
[0112] Another embodiment of the vehicle path planning method provided by the present application may specifically include the following steps:
[0113] S21, during the driving process of the vehicle, obtaining an obstacle vehicle corresponding to the current vehicle.
[0114] S22, taking the root node as a parent node and generating at least one child node.
[0115] S23, taking the child node as the parent node, and continuing to execute the step of generating at least one child node until the number of layers of the Monte Carlo search tree is greater than or equal to the preset number of layers.
[0116] Among them, the first state includes the current acceleration of the current vehicle, the first target state includes the target acceleration, the first target heading angle, and the first target curvature of the current vehicle, the score includes the driving comfort, and the score includes the driving comfort.
[0117] S24, when the number of layers of the Monte Carlo search tree is greater than or equal to the preset number of layers, the acceleration comfort corresponding to the current node is obtained by using the current acceleration and the target acceleration.
[0118] In some embodiments, when the number of layers of the Monte Carlo search tree is greater than or equal to a preset number of layers, the acceleration comfort corresponding to the current node is obtained using the current acceleration and the target acceleration of the current vehicle corresponding to the current node.
[0119] In some embodiments, the acceleration comfort satisfies the following relationship:
[0120] jerk_smoothness_cost=k_smooth*(k1*abs(acc_ca)+abs(k2*jerk_smoothness)+k3*abs(car_state.curret_accl))
[0121] Among them, jerk_smoothness_cost is the acceleration comfort, acc_ca is the current acceleration of the current vehicle at the current moment, car_state.curret_accl is the target acceleration corresponding to the target node, and jerk_smoothness is the difference between the target acceleration and the current acceleration.
[0122] The value of k_smooth is determined by whether the direction of the target acceleration and the jerk (i.e., the derivative of the acceleration) are consistent. If they are consistent, k_smooth is 1.6. If they are inconsistent, k_smooth is 0.6. Furthermore, k1 is 10, k2 is 100, and k3 is 10.
[0123] S25, using the first target heading angle and the second target heading angle of the parent node corresponding to the current node, obtain the heading angle comfort corresponding to the current node.
[0124] In some embodiments, the heading angle comfort satisfies the following relationship:
[0125] theta_smoothness_cost=1.0 / (1.0+k_theta_smoothness*exp(theta_smoothness_deviation-theta_smoothness))
[0126] Among them, theta_smoothness_cost is the heading angle comfort, k_theta_smoothness is a coefficient, which can be 1.5, kappa_smoothness_deviation is the first target heading angle, and kappa_smoothness is the second target heading angle.
[0127] S26, using the first target curvature and the second target curvature of the parent node, obtain the curvature comfort corresponding to the current node.
[0128] In some embodiments, the first target curvature can be determined based on the relative position between the current node and the parent node, and the relative vehicle head orientation between the current node and the parent node. The second target curvature can also be determined in the above manner.
[0129] In some embodiments, the curvature comfort satisfies the following relationship:
[0130] kappa_smoothness_cost=1.0 / (1.0+k_kappa_smoothness*exp(kappa_smoothness_deviatioin-kappa_smoothness))
[0131] Among them, kappa_smoothness_cost is the curvature comfort, k_kappa_smoothness is the coefficient, which can be 1.5, kappa_smoothness_deviation is the first target curvature, and kappa_smoothness is the second target curvature.
[0132] S27, calculating the weighted sum of the acceleration comfort, the heading angle comfort, and the curvature comfort to obtain the driving comfort of the current node.
[0133] In some embodiments, the driving comfort satisfies the following relationship:
[0134] smoothness_cost=u1*jerk_smoothness_cost+u2*theta_smoothness_cost+u3*kappa_smoothness_cost
[0135] Among them, smoothness_cost is the driving comfort, and u1, u2, and u3 are weight coefficients respectively.
[0136] S28, using the score to determine the UCT value corresponding to each node.
[0137] In some embodiments, the UCT value of a certain node m satisfies the following relationship:
[0138]
[0139] Where Q(m) is the average score of node m, N is the number of visits to the parent node of node m, n is the number of visits to node m, and c is a constant that controls the balance between exploration and exploitation.
[0140] In some embodiments, Q(m) is the quotient of the score of node m and the number of visits to node m, that is, it satisfies the following relationship:
[0141]
[0142] Among them, cost_m is the score of node m.
[0143] S29, determining a plurality of decision branches according to the UCT value corresponding to each node in the Monte Carlo search tree, and determining the best decision branch from the plurality of decision branches.
[0144] S30, controlling the current vehicle driving according to the best decision branch.
[0145] In the above scheme, during the driving process of the vehicle, the obstacle vehicle corresponding to the current vehicle is obtained, and a Monte Carlo search tree is constructed based on the first state of the current vehicle and the second state of the obstacle vehicle; a number of decision branches are determined according to the UCT value corresponding to each node in the Monte Carlo search tree, and then the driving of the current vehicle is controlled according to the best decision branch. The path planning method provided by the present application uses the state of the current vehicle at the current moment as the root node, performs a Monte Carlo tree search on the future running state of the current vehicle, and adds the driving state with the highest UCT value to the best decision branch. The best decision branch obtained can reduce the possibility of collision between the current vehicle and the obstacle vehicle and improve the driving safety of the current vehicle.
[0146] Another embodiment of the vehicle path planning method provided by the present application may specifically include the following steps:
[0147] S31, during the driving process of the vehicle, obtaining an obstacle vehicle corresponding to the current vehicle.
[0148] S32, taking the root node as a parent node and generating at least one child node.
[0149] S33, taking the child node as the parent node, and continuing to execute the step of generating at least one child node until the number of layers of the Monte Carlo search tree is greater than or equal to the preset number of layers.
[0150] The first target state includes the target speed and the first target position, and the score includes the traffic efficiency.
[0151] S34, using the target speed and the maximum speed limit of the road where the vehicle is currently located, obtain the speed efficiency corresponding to the current node.
[0152] In some embodiments, the speed efficiency satisfies the following relationship:
[0153] speed_efficiency_cost=fmax(speed_efficiency_cost_o,0)
[0154] Furthermore, speed_efficiency_cost_o satisfies the following relationship:
[0155] speed_efficiency_cost_o=2.9 / (curret_speed / speedlimit)
[0156] Among them, current_speed is the current target (longitudinal) speed of the vehicle, and speedlimit is the maximum speed limit.
[0157] S35, obtaining all associated nodes between the root node and the current node, and determining the target path length using the position information corresponding to all associated nodes.
[0158] Exemplarily, the node layer where the current node is located is the 4th layer, and the path planning device obtains the root node, the 1st layer node corresponding to the current node, the 2nd layer node corresponding to the current node, and the 3rd layer node corresponding to the current node. Further, the 1st layer node, the 2nd layer node, and the 3rd layer node are associated nodes.
[0159] Further, the path planning device determines the driving path of the current vehicle before the current node, that is, the target path, based on the location information of the root node and the location information corresponding to all associated nodes. Further, the path planning device calculates the length of the target path, that is, the target path length path_length.
[0160] S36, using the target longitudinal distance and the target path length, obtain the location traffic efficiency corresponding to the current node.
[0161] The target longitudinal distance is the longitudinal distance between the first target position and the second target position of the parent node corresponding to the current node.
[0162] In some embodiments, the location traffic efficiency satisfies the following relationship:
[0163] travel_efficiency_cost=fmax(k_quad*(tar get_distance-distance_s),0)
[0164] Among them, travel_efficiency_cost is the location travel efficiency, and distance_s is the target longitudinal distance.
[0165] Furthermore, k_quad satisfies the following relationship:
[0166] k_quad = 1 / distance_s
[0167] Furthermore, target_distance satisfies the following relationship:
[0168] target_distance=path_length-ego_s
[0169] Among them, path_length is the target path length, and ego_s is the longitudinal position of the current vehicle at the current node in the Frenet coordinate system.
[0170] S37, calculating the weighted sum of the speed efficiency and the position efficiency to obtain the efficiency of the current node.
[0171] In some embodiments, the traffic efficiency satisfies the following relationship:
[0172] efficiency_cost=u4*speed_efficiency_cost+u5*travel_efficiency_cost
[0173] Among them, efficiency_cost is the traffic efficiency, and u4 and u5 are weight values respectively.
[0174] S38, using the score to determine the UCT value corresponding to each node.
[0175] S39, determining a plurality of decision branches according to the UCT value corresponding to each node in the Monte Carlo search tree, and determining the best decision branch from the plurality of decision branches.
[0176] S40, controlling the current vehicle driving according to the best decision branch.
[0177] In the above scheme, during the driving process of the vehicle, the obstacle vehicle corresponding to the current vehicle is obtained, and a Monte Carlo search tree is constructed based on the first state of the current vehicle and the second state of the obstacle vehicle; a number of decision branches are determined according to the UCT value corresponding to each node in the Monte Carlo search tree, and then the driving of the current vehicle is controlled according to the best decision branch. The path planning method provided by the present application uses the state of the current vehicle at the current moment as the root node, performs a Monte Carlo tree search on the future running state of the current vehicle, and adds the driving state with the highest UCT value to the best decision branch. The best decision branch obtained can reduce the possibility of collision between the current vehicle and the obstacle vehicle and improve the driving safety of the current vehicle.
[0178] Another embodiment of the vehicle path planning method provided by the present application may specifically include the following steps:
[0179] S41, during the driving process of the vehicle, obtaining an obstacle vehicle corresponding to the current vehicle.
[0180] S42, taking the root node as a parent node and generating at least one child node.
[0181] S43, taking the child node as the parent node, and continuing to execute the step of generating at least one child node until the number of layers of the Monte Carlo search tree is greater than or equal to the preset number of layers.
[0182] The first target state includes the target acceleration, and the score includes the historical score.
[0183] S44, obtaining the historical acceleration of the current vehicle at the historical moment, and determining a preset acceleration range according to the historical acceleration.
[0184] In some embodiments, the historical moment may be any one of 0.1s-1s before the current moment.
[0185] Exemplarily, the historical moments may include 0.1s, 0.2s, 0.3s, 0.4s, 0.5s, and 0.6s before the current moment. Further, the path planning device obtains the accelerations 0.1s, 0.2s, 0.3s, 0.4s, 0.5s, and 0.6s before the current moment to obtain a historical acceleration set.
[0186] In some embodiments, the path planning device determines a preset acceleration interval based on the historical acceleration. For example, the minimum value of the preset acceleration interval may be 0.8 times, 0.85 times, or 0.9 times the historical acceleration, and the maximum value of the preset acceleration interval may be 1.1 times, 1.15 times, or 1.2 times the historical acceleration.
[0187] S45: In response to the target acceleration being within the preset acceleration interval, adding a first preset value to the historical score.
[0188] In some embodiments, in response to the target acceleration being within a preset acceleration interval, the path planning device adds a first preset value to the third score, wherein the first preset value is a positive number.
[0189] In some possible embodiments, the historical moment is 0.1s before the current moment, and the corresponding first preset value is 1000. In some possible embodiments, the historical moment is 0.2s before the current moment, and the corresponding first preset value is 800. In some possible embodiments, the historical moment is 0.3s before the current moment, and the corresponding first preset value is 400. In some possible embodiments, the historical moment is 0.4s before the current moment, and the corresponding first preset value is 300. In some possible embodiments, the historical moment is 0.5s before the current moment, and the corresponding first preset value is 200. In some possible embodiments, the historical moment is 0.6s before the current moment, and the corresponding first preset value is 100.
[0190] In some embodiments, in response to the target acceleration being outside the preset acceleration range, the path planning device does not add the first preset value to the third score, wherein the initial value of the third score is 0.
[0191] S46, using the score to determine the UCT value corresponding to each node.
[0192] S47, determining a plurality of decision branches according to the UCT value corresponding to each node in the Monte Carlo search tree, and determining the best decision branch from the plurality of decision branches.
[0193] S48, controlling the current vehicle driving according to the best decision branch.
[0194] In the above scheme, during the driving process of the vehicle, the obstacle vehicle corresponding to the current vehicle is obtained, and a Monte Carlo search tree is constructed based on the first state of the current vehicle and the second state of the obstacle vehicle; a number of decision branches are determined according to the UCT value corresponding to each node in the Monte Carlo search tree, and then the driving of the current vehicle is controlled according to the best decision branch. The path planning method provided by the present application uses the state of the current vehicle at the current moment as the root node, performs a Monte Carlo tree search on the future running state of the current vehicle, and adds the driving state with the highest UCT value to the best decision branch. The best decision branch obtained can reduce the possibility of collision between the current vehicle and the obstacle vehicle and improve the driving safety of the current vehicle.
[0195] Another embodiment of the vehicle path planning method provided by the present application may specifically include the following steps:
[0196] S51, during the driving process of the vehicle, obtaining an obstacle vehicle corresponding to the current vehicle.
[0197] S52: Take the root node as the parent node and generate at least one child node.
[0198] S53, taking the child node as the parent node, and continuing to execute the step of generating at least one child node until the number of layers of the Monte Carlo search tree is greater than or equal to the preset number of layers.
[0199] Among them, the score includes lane change penalty items.
[0200] S54, judging whether the current vehicle changes lanes according to the second target position of the target node corresponding to the parent node and the first target position.
[0201] Specifically, the path planning device obtains the second lateral position in the second target position and the first lateral position in the first target position, and determines whether the current vehicle changes lanes by calculating the difference between the first lateral position and the second lateral position.
[0202] In some embodiments, the path planning device calculates the absolute value of the difference between the first lateral position and the second lateral position, and if the absolute value is greater than or equal to a preset threshold, it is determined that the current vehicle has changed lanes.
[0203] S55: In response to the current vehicle changing lanes, adding a second preset value to the lane-changing penalty item.
[0204] Specifically, in response to the current lane-changing behavior of the vehicle, the path planning device adds the second preset value to the lane-changing penalty item.
[0205] In some embodiments, the lane changing penalty term is represented by changing_lane_cost.
[0206] In some embodiments, the second preset value is -3000.
[0207] In some embodiments, in response to the current vehicle not having a lane change behavior, the path planning device does not add the second preset value to the lane change penalty item, wherein the initial value of the lane change penalty item is 0.
[0208] By adding a lane change penalty item, the lateral smoothness of the current vehicle is improved.
[0209] S56, using the score to determine the UCT value corresponding to each node.
[0210] S57, determining a plurality of decision branches according to the UCT value corresponding to each node in the Monte Carlo search tree, and determining the best decision branch from the plurality of decision branches.
[0211] S58, controlling the current vehicle driving according to the best decision branch.
[0212] In the above scheme, during the driving process of the vehicle, the obstacle vehicle corresponding to the current vehicle is obtained, and a Monte Carlo search tree is constructed based on the first state of the current vehicle and the second state of the obstacle vehicle; a number of decision branches are determined according to the UCT value corresponding to each node in the Monte Carlo search tree, and then the driving of the current vehicle is controlled according to the best decision branch. The path planning method provided by the present application uses the state of the current vehicle at the current moment as the root node, performs a Monte Carlo tree search on the future running state of the current vehicle, and adds the driving state with the highest UCT value to the best decision branch. The best decision branch obtained can reduce the possibility of collision between the current vehicle and the obstacle vehicle and improve the driving safety of the current vehicle.
[0213] Another embodiment of the vehicle path planning method provided by the present application may specifically include the following steps:
[0214] S61, during the driving process of the vehicle, obtaining an obstacle vehicle corresponding to the current vehicle.
[0215] S62: Take the root node as the parent node and generate at least one child node.
[0216] The subnode is used to simulate the first target state of the current vehicle after the first preset time period.
[0217] In some embodiments, the path planning device also updates the second target state of the obstructing vehicle after the first preset time period according to the second state corresponding to the obstructing vehicle. For example, the path planning device updates the position and speed of the obstructing vehicle after the first preset time period according to the current position, current speed, current acceleration, and current vehicle head direction corresponding to the obstructing vehicle.
[0218] Wherein, the current vehicle in the first target state does not collide with the obstacle vehicle in the second target state;
[0219] S63, taking the child node as the parent node, and continuing to execute the step of generating at least one child node until the number of layers of the Monte Carlo search tree is greater than or equal to the preset number of layers.
[0220] S64, when the number of layers of the Monte Carlo search tree is greater than or equal to the preset number of layers, the score corresponding to each node is calculated, and the UCT value corresponding to each node is determined using the score.
[0221] S65, using the root node, construct a decision branch.
[0222] Specifically, the path planning device uses the root node of the Monte Carlo tree to construct a decision branch.
[0223] S66, taking the root node as the parent node, and adding the target child node with the largest UCT value from the child nodes of the next layer to the decision branch.
[0224] Specifically, the path planning device takes the root node as the parent node, and adds the target child node with the largest UCT value from the child nodes of the next layer to the decision branch.
[0225] It should be noted that the target subnode with the largest UCT value indicates that its corresponding current vehicle operating state is more in line with the requirements.
[0226] S67, taking the target child node as the parent node, repeatedly adding the target child node with the largest UCT value from the next layer of child nodes to the decision branch until the number of layers where the target child node is located is greater than or equal to the preset number of layers.
[0227] Specifically, the path planning device takes the target subnode as the parent node, and repeatedly adds the target subnode with the largest UCT value from the next layer of subnodes to the decision branch until the number of layers where the target subnode is located is greater than or equal to the preset number of layers.
[0228] S68, controlling the current vehicle driving according to the best decision branch.
[0229] In the above scheme, during the driving process of the vehicle, the obstacle vehicle corresponding to the current vehicle is obtained, and a Monte Carlo search tree is constructed based on the first state of the current vehicle and the second state of the obstacle vehicle; a number of decision branches are determined according to the UCT value corresponding to each node in the Monte Carlo search tree, and then the driving of the current vehicle is controlled according to the best decision branch. The path planning method provided by the present application uses the state of the current vehicle at the current moment as the root node, performs a Monte Carlo tree search on the future running state of the current vehicle, and adds the driving state with the highest UCT value to the best decision branch. The best decision branch obtained can reduce the possibility of collision between the current vehicle and the obstacle vehicle and improve the driving safety of the current vehicle.
[0230] Another embodiment of the vehicle path planning method provided by the present application may specifically include the following steps:
[0231] S71, obtaining candidate vehicles around the current vehicle.
[0232] Specifically, the path planning device obtains candidate vehicles around the current vehicle, and obtains information such as the position, speed, and front direction of the candidate vehicles.
[0233] In this embodiment, "surroundings" refers to other vehicles whose lateral distance from the current vehicle is less than the lateral distance threshold, and / or whose longitudinal distance from the current vehicle is less than the longitudinal distance threshold. It can be understood that the longitudinal distance threshold can be determined based on the longitudinal speed of other vehicles and / or the current vehicle. For example, when the speed of other vehicles and / or the current vehicle is relatively high, it corresponds to the first longitudinal distance threshold, and when the speed of other vehicles and / or the current vehicle is relatively low, it corresponds to the second longitudinal distance threshold, that is, the first longitudinal distance threshold is greater than the second longitudinal distance threshold. Similarly, the lateral distance threshold can be determined based on the longitudinal speed of other vehicles and / or the current vehicle.
[0234] In some embodiments, the path planning device further obtains the size information of the candidate vehicle, which may include the maximum length, maximum width and maximum height of the candidate vehicle. Further, the path planning device uses this to construct a first maximum circumscribed cuboid corresponding to the candidate vehicle. It can be understood that the length of the first maximum cuboid is the maximum length of the candidate vehicle, the width of the first maximum cuboid is the maximum width of the candidate vehicle, and the height of the first maximum cuboid is the maximum height of the candidate vehicle.
[0235] Similarly, the path planning device constructs the second largest circumscribed cuboid corresponding to the current vehicle according to the maximum length, maximum width and maximum height of the current vehicle. Similarly, the length of the second largest cuboid is the maximum length of the current vehicle, the width of the second largest cuboid is the maximum width of the current vehicle, and the height of the second largest cuboid is the maximum height of the current vehicle.
[0236] S72: In response to a collision between the candidate vehicle and the current vehicle within a second preset time period, treating the candidate vehicle as an obstacle vehicle.
[0237] In some embodiments, the path planning device determines the first driving path of the current vehicle within the second preset time period based on the current position, current speed, current vehicle head direction, and current acceleration of the current vehicle. Similarly, the path planning device determines the second driving path of the candidate vehicle within the second preset time period based on the current position, current speed, current vehicle head direction, and current acceleration of the candidate vehicle. Furthermore, the path planning device determines whether the first largest circumscribed cuboid representing the current vehicle and the second largest circumscribed cuboid representing the candidate vehicle will be in contact based on the first driving path and the second driving path. If they are in contact, it means that the candidate vehicle collides with the current vehicle, and the path planning device will treat the candidate vehicle that meets the above conditions as an obstacle vehicle.
[0238] S73, constructing a Monte Carlo search tree according to the first state of the current vehicle and the second state of the obstacle vehicle.
[0239] Among them, the first state includes the current position, current speed and current front direction of the current vehicle at the current moment; the second state includes the current position, current speed and current front direction of the obstacle vehicle; among them, the first state is the root node of the Monte Carlo search tree.
[0240] S74, determining a plurality of decision branches according to the UCT value corresponding to each node in the Monte Carlo search tree, and determining the best decision branch from the plurality of decision branches.
[0241] Among them, the node is used to simulate the operating status of the current vehicle at a future moment.
[0242] S75, controlling the current vehicle driving according to the best decision branch.
[0243] In the above scheme, during the driving process of the vehicle, the obstacle vehicle corresponding to the current vehicle is obtained, and a Monte Carlo search tree is constructed based on the first state of the current vehicle and the second state of the obstacle vehicle; a number of decision branches are determined according to the UCT value corresponding to each node in the Monte Carlo search tree, and then the driving of the current vehicle is controlled according to the best decision branch. The path planning method provided by the present application uses the state of the current vehicle at the current moment as the root node, performs a Monte Carlo tree search on the future running state of the current vehicle, and adds the driving state with the highest UCT value to the best decision branch. The best decision branch obtained can reduce the possibility of collision between the current vehicle and the obstacle vehicle and improve the driving safety of the current vehicle.
[0244] See also Figure 2 , Figure 2 It is a flow chart of another embodiment of the vehicle path planning method provided in the present application.
[0245] like Figure 2 As shown, another embodiment of the vehicle path planning method provided by the present application may specifically include the following steps:
[0246] S101, obtaining the vehicle speed and position information, the speed and position information of surrounding vehicles, and the lane line information.
[0247] Among them, the ego vehicle corresponds to the current vehicle mentioned above, and the surrounding vehicles correspond to the candidate vehicles mentioned above.
[0248] In some embodiments, the path planning device obtains the current state of the vehicle, for example, obtains speed information, acceleration information, positioning information and other related information from the chassis.
[0249] At the same time, the path planning device uses the sensors carried by the vehicle to obtain information about the surrounding environment. The surrounding environment information mainly includes status information (speed information, acceleration information, vehicle head direction) and position information of surrounding vehicles.
[0250] Specifically, the information obtained by the path planning device is shown in the following table.
[0251] Self-driving status (x,y,s,l,speed,heading,accel,size) Front vehicle status (x,y,speed,heading,accel,size) Rear vehicle status (x,y,speed,heading,accel,size) Self-driving reference line {(x,y,s,heading,speed_limit),...} Predicted trajectory of the preceding vehicle {(x,y,speed,heading),...}
[0252] Among them, x refers to the horizontal coordinate (lateral position) in the global coordinate system, y refers to the vertical coordinate (longitudinal position) in the global coordinate system, s refers to the horizontal coordinate (lateral position) in the Frenet coordinate system, l refers to the vertical coordinate (longitudinal position) in the Frenet coordinate system, speed refers to speed, heading refers to the direction of the vehicle head, accel refers to acceleration, and size refers to the vehicle size.
[0253] In some embodiments, the path planning device uses sensors carried by the vehicle to obtain lane information.
[0254] After obtaining information, the surrounding obstacle vehicles are screened, and vehicles with strong interactive correlation are screened. The screening conditions mainly consider whether the position of the surrounding obstacle vehicles is close to the current position of the vehicle, and whether there is overlap in a short period of time in the future.
[0255] S102, determining whether surrounding vehicles are interactive vehicles.
[0256] If the judgment result is yes, jump to S103. If the judgment result is no, jump to S104.
[0257] Among them, the interactive vehicle corresponds to the obstacle vehicle mentioned above.
[0258] In some embodiments, the path planning device determines whether the surrounding vehicles are interactive vehicles based on whether the relative positions between the surrounding vehicles and the own vehicle are small, and whether the surrounding vehicles and the own vehicle overlap (overlap means collision) in the future.
[0259] See also Figure 3 , the ego vehicle is a vehicle merging into the main road. At this time, there are interactive vehicles 1 and 2 on the main road. Interactive vehicle 1 is a forward obstacle vehicle, and interactive vehicle 2 is a backward obstacle vehicle. That is, any state change made by either party (interactive vehicle or ego vehicle) will affect the action choice of the other party.
[0260] S103, performing path planning according to the IDM method.
[0261] Specifically, the path planning device plans the path of the vehicle according to IDM (Intelligent Driver Model).
[0262] In some embodiments, the path planning device performs lane change and acceleration / deceleration planning based on the relative position and relative speed of the vehicle in front of the vehicle.
[0263] S104, using the Monte Carlo search tree to perform decision planning and obtain the best decision branch.
[0264] In some embodiments, in response to the absence of an interacting vehicle, the path planning device directly selects the most comfortable lane change under the conditions of the road environment or selects a default lane change rule to plan a lane change trajectory.
[0265] In some embodiments, in response to the existence of interacting vehicles, the path planning device uses a Monte Carlo search tree to perform decision planning to obtain an optimal decision branch.
[0266] The Monte Carlo search tree explores different decision paths by building a search tree and evaluates the value of these paths through simulation. A complete Monte Carlo search tree mainly consists of four stages: selection, expansion, simulation, and backpropagation.
[0267] In some embodiments, the Monte Carlo search tree uses the UCT value of each node to select the best decision branch. Specifically, the UCT formula satisfies the following relationship:
[0268]
[0269] Where Q(m) is the average score of node m, N is the number of visits to the parent node of node m, n is the number of visits to node m, and c is a constant that controls the balance between exploration and exploitation.
[0270] The selection process continues until it reaches an incompletely expanded node or leaf node. Once an incompletely expanded node or leaf node is reached, the Monte Carlo search tree will create one or more child nodes. These child nodes represent possible actions starting from the current node. If all possible actions have been explored, it goes directly to the simulation phase. In the simulation phase, the Monte Carlo search tree starts from the newly created node and uses some heuristic method (usually a random strategy) to simulate a complete sequence until it reaches the terminal state of the game or reaches a preset depth limit. This process is called simulation or rollout.
[0271] After the simulation is complete, the statistics of all nodes on the path from the root node to the simulated node are updated. This usually involves updating the node's visit count and accumulated reward.
[0272] In the embodiment of the present application, the definition of a node is the state of the vehicle. The root node root is defined as the current state information of the vehicle. The expansion node (child node) is the deduced state of the vehicle. In addition, each node should also include the score of the current node, the average score, the UCT value, the index of the node in the entire Monte Carlo search tree, the number of visits, the number of node layers, the child node information, and the parent node information.
[0273] The current node score is composed of multiple parts, including driving comfort, traffic efficiency, terminal traffic efficiency, lane change rewards, safety penalties, and historical results. The total score of node m is the sum of the rewards calculated after each visit to node m, and the average score of node m is the total score divided by the number of visits. The UCT value is obtained according to the UCT calculation formula above. The node index is the index in the MCTS decision tree. The number of node layers is the number of deduction steps of the current node. The deduction time of the current node (the length of time from the node moment to the current moment) can be determined based on the number of node layers and the deduction step length. The child node information is the node expanded based on the current node. The parent node is opposite to the child node.
[0274] Among them, the current node score corresponds to the score above.
[0275] The current node score specifically satisfies the following relationship:
[0276] cost_all=smoothness_cost+efficiency_cost+reliability_cost+lane_select_cost+safety_cost+history_cost
[0277] Among them, cost_all is the current node score, smoothness_cost is the driving comfort, efficiency_cost is the traffic efficiency, reliability_cost is the location traffic efficiency, lane_select_cost is the lane change penalty, safety_cost is the safety penalty, and history_cost is the history penalty.
[0278] Among them, the driving comfort satisfies the following relationship:
[0279] smoothness_cost=jerk_smoothness_cost+theta_smoothness_cost+kappa_smoothness_cost
[0280] Among them, jerk_smoothness_cost is the longitudinal comfort, theta_smoothness_cost is the lateral comfort, and kappa_smoothness_cost is the curvature comfort.
[0281] Furthermore, the longitudinal comfort satisfies the following relationship:
[0282] jerk_smoothness_cost=k_smooth*(k1*abs(acc_ca)+abs(k2*jerk_smoothness_)+k3*abs(curret_accl))
[0283] Among them, k_smooth is determined by whether the acceleration and jerk direction of the node are consistent. If they are consistent, the value is 1.5, if not, the value is 0.6. The sub-weight coefficient k1 is 10, k2 is 100, and k3 is 10. acc_ca is the acceleration corresponding to the current vehicle state at the current moment. jerk_smoothness_ is the difference between the acceleration corresponding to the current node and acc_ca. curret_accl is the acceleration corresponding to the current node.
[0284] Among them, the lateral comfort satisfies the following relationship:
[0285] theta_smoothness_cost=1.0 / (1.0+k_theta_smoothness*exp(theta_smoothness_deviation-theta_smoothness))
[0286] Among them, k_theta_smoothness is a coefficient term, which can be 1.5, theta_smoothness_deviation is the heading angle of the current node, and theta_smoothness is the Hagen inverse angle of the parent node corresponding to the current node.
[0287] Among them, the curvature comfort satisfies the following relationship:
[0288] kappa_smoothness_cost=1.0 / (1.0+k_kappa_smoothness*exp(kappa_smoothness_deviation-kappa_smoothness))
[0289] Among them, k_kappa_smoothness is a coefficient term, which can be 1.5, kappa_smoothness_deviation is the curvature corresponding to the current node, and kappa_smoothness is the curvature of the parent node corresponding to the current node.
[0290] The path planning device may calculate the curvature according to the horizontal coordinate, the vertical coordinate and the vehicle head direction of the current node and the parent node corresponding to the current node in the global coordinate system.
[0291] Specifically, the traffic efficiency satisfies the following relationship:
[0292] efficiency_cost=speed_efficiency_cost+travel_efficiency_cost
[0293] Among them, speed_efficiency_cost is the speed efficiency, and travel_efficiency_cost is the location efficiency.
[0294] Among them, the speed efficiency satisfies the following relationship:
[0295] speed_efficiency_cost=fmax(speed_efficiency_cost,0);
[0296] in:
[0297] speed_efficiency_cost=2.9 / (curret_speed / speedlimit)
[0298] speedlimit is the maximum speed limit of the current road, and curret_speed is the speed of the vehicle at the current node.
[0299] The location efficiency satisfies the following relationship:
[0300] avel_efficiency_cost=fmax(k_quad*(target_distance-distance_s),0)
[0301] Where: k_quad = 1 / distance_s, distance_s is the longitudinal movement distance of the vehicle
[0302] target_distance satisfies the following relationship:
[0303] target_distance=path_lenngth-ego_s
[0304] path_lenngth is the total path length generated from the root node to the parent node corresponding to the current node, and ego_s is the longitudinal position of the current vehicle in the Frenet coordinate system corresponding to the current node.
[0305] The location efficiency satisfies the following relationship:
[0306] reliability_cost=deviation_cost
[0307] Deviation_cost is the penalty term for the distance from the target point. If the vehicle reaches within 10m of the navigation target point, a negative penalty is added with a penalty value of -1000, otherwise it is 0.
[0308] The lane change penalty satisfies the following relationship:
[0309] lane_select_cost=changing_lane_cost
[0310] In order to increase lateral smoothness, a penalty is imposed on lane changing behavior. If the ego vehicle changes lanes, the changing_lane_cost value is -3000, otherwise it is 0.
[0311] The security penalty satisfies the following relationship:
[0312] safety_cost=collision_cost+nervous_cost
[0313] Among them, collision_cost is the collision penalty and nervous_cost is the safety reward.
[0314] In some embodiments, the path planning device determines whether there is overlap (overlap means collision) between the three-dimensional model corresponding to the own vehicle (not the maximum circumscribed cuboid mentioned above) and the three-dimensional model corresponding to the interacting vehicle. The more overlap there is, the smaller the collision penalty is. The minimum value of the collision penalty is -5000.
[0315] In some embodiments, the path planning device calculates the relative distance between the self-vehicle and the interactive vehicle. The larger the relative distance, the larger the safety distance. The maximum value of the safety distance is 1500.
[0316] Specifically, the historical results satisfy the following relationship:
[0317] history_cost=last_result_cost+history_decision_result
[0318] Among them, history_cost is the historical result. last_result_cost is the result of one frame before the current moment. history_decision_result is the result of 2-6 frames before the current moment. Among them, 1 frame refers to 0.1s.
[0319] The cost of the benchmark historical cost item is 1000. The cost coefficients of historical frames 1-6 are {1, 0.8, 0.4, 0.3, 0.2, 0.1}. Compare the result of the current node with the acceleration of historical frames 1-6. If the error is within 0.1, it is considered consistent. If consistent, a positive value is applied to the current node.
[0320] In some embodiments, the process of building a node may include the following steps:
[0321] First, initialize the node's score cost to 0 and set the current node currNode to the input node.
[0322] Furthermore, the path planning device performs random deduction within a certain range and generates a random number randomNum. The range of the random number randomNum is between 0 and 1. And based on the random number, a random acceleration randomAcc is generated. According to the value of randomNum, one of the following three behaviors is selected: If randomNum<=0.1, decelerate. If randomNum>=0.9, accelerate. If the above conditions are not met, maintain the speed. It should be noted that the range of acceleration values of the vehicle can be different in different road scenarios. For example, in the highway scenario, the acceleration value is between -3m / s 2 Up to 3m / s 2 In the normal road scenario, the acceleration value is between -2m / s 2 Up to 2m / s 2 between.
[0323] Furthermore, the path planning device calculates the change in acceleration (deltaAcc) and the rate of change of acceleration (Jerk). Among them, the rate of change of acceleration corresponds to the jerk mentioned above. The path planning calls the getDisplacement function to obtain the displacement (displacementS, displacementL) of the vehicle in the Frenet coordinate system and the speed change (deltaSpeedS, deltaSpeedL) in the Frenet coordinate system. Furthermore, the path planning device creates a new child node newNode under the current node and updates its position state, time, Frenet position state and other attributes. And call the Frenet2global function to convert the vehicle's position state under Frenet to the position state in the global coordinate system. Call the checkCollision function to check whether the vehicle and the interactive vehicle collide.
[0324] In some embodiments, when the ego vehicle collides with the interactive vehicle, the path planning device sets cost_all of the node to -10000 and exits the loop.
[0325] In some embodiments, the ego vehicle and the interactive vehicle do not collide, update the current node currNode to newNode and continue the loop. This function (checkCollision function) evaluates the future cost of an MCTS node by simulating different behaviors of the vehicle (deceleration, acceleration, maintaining speed) and checking whether a collision occurs at each step. If no collision occurs, the simulation continues until the maximum time range is reached. If a collision occurs, a large negative penalty value is returned to indicate that an unsafe situation has occurred.
[0326] expandDetailed method: Generate new nodes by simulating different vehicle behaviors (such as emergency braking, maintaining a constant speed, changing lanes, accelerating / decelerating), and check whether these behaviors lead to a collision. If no collision occurs, a new node is added to the tree. This allows multiple possible behavior paths to be explored and provides richer options for subsequent simulations and selections. The different behaviors of the current vehicle are specifically calculated as follows:
[0327] (1) Emergency braking. Calculate the emergency braking acceleration emergencyAcc. If emergencyAcc is less than -8m / s 2 , then limit it to -8m / s 2. Calculate the emergency braking acceleration change rate emergencyJerkS. Call the getDisplacement function to calculate the displacement and velocity change. Create a new node newNode5 and update its Frenet state, global state and other properties. If no collision occurs and the current velocity is greater than 0, add the new node to the tree.
[0328] (2) Maintain a constant speed. Create a new node newNode3 and calculate the acceleration change rate Jerk3. Call the getDisplacement function to calculate the displacement and velocity change. Update the Frenet state, global state and other properties of the new node newNode3. If no collision occurs and the speed is within the speed limit, add the new node to the tree.
[0329] (3) Lane change. Traverse the roadside boundary lbdry to determine whether the current vehicle can change lanes. If it can change lanes, create a new node newNode4 and update its Frenet state, global state and other attributes. If there is no collision and the speed is within the speed limit, add the new node to the tree.
[0330] (4) Acceleration / deceleration. Traverse all possible values of accMax and create new nodes for acceleration / deceleration. For deceleration, if the current speed is greater than 1m / s, create a new node. For acceleration, create a new node and update its Frenet state, global state and other attributes. If no collision occurs and other conditions are met, add the new node to the tree.
[0331] In some embodiments, when the Monte Carlo search tree stops iterating, the path planning device starts from the root node, selects the node with the highest UCT value from each layer of nodes, and adds it to the best decision branch.
[0332] S105, using the best decision branch to control the current vehicle driving.
[0333] In the above scheme, during the driving process of the vehicle, the obstacle vehicle corresponding to the current vehicle is obtained, and a Monte Carlo search tree is constructed based on the first state of the current vehicle and the second state of the obstacle vehicle; a number of decision branches are determined according to the UCT value corresponding to each node in the Monte Carlo search tree, and then the driving of the current vehicle is controlled according to the best decision branch. The path planning method provided by the present application uses the state of the current vehicle at the current moment as the root node, performs a Monte Carlo tree search on the future running state of the current vehicle, and adds the driving state with the highest UCT value to the best decision branch. The best decision branch obtained can reduce the possibility of collision between the current vehicle and the obstacle vehicle and improve the driving safety of the current vehicle.
[0334] Furthermore, in this embodiment, the Monte Carlo search tree algorithm is used to expand the search tree, which means that future possibilities can be explored by simulating different behavior paths. Since the Monte Carlo search tree is a heuristic search algorithm, it tends to explore the most promising path rather than blindly trying all possible paths. This enables the algorithm to find a relatively good solution within a limited time. This approach improves decision-making efficiency because high-quality behavior paths can be found faster without considering all possible behavior combinations.
[0335] The embodiment of the present application calls the checkCollision function to check whether a collision will occur before creating a new node. This means that only safe behavior paths will be added to the search tree. This collision detection mechanism significantly enhances the security of the system because it can identify and exclude behavior paths that may cause collisions in advance.
[0336] The embodiments of the present application can handle a variety of vehicle behaviors, including emergency braking, maintaining a constant speed, changing lanes, accelerating / decelerating, etc. This means that the algorithm can adapt to different driving scenarios and emergency situations, so as to make appropriate decisions in various complex traffic environments, thereby improving overall adaptability and robustness.
[0337] The embodiments of the present application take into account multiple factors when creating new nodes, such as speed limits, acceleration limits, etc. This means that multiple objectives (such as safety, efficiency, comfort, etc.) can be optimized at the same time. This multi-objective optimization capability can balance different objectives and ultimately produce a decision-making solution that is both safe and efficient.
[0338] Please continue to see Figure 4 , Figure 4 The terminal device 500 of the embodiment of the present application includes a processor 51 and a memory 52 .
[0339] The processor 51 and the memory 52 are connected to the bus. The memory 52 stores program data. The processor 51 is used to execute the program data to implement the vehicle path planning method described in the above embodiment.
[0340] In the embodiment of the present application, the processor 51 may also be referred to as a CPU (Central Processing Unit). The processor 51 may be an integrated circuit chip having the ability to process signals. The processor 51 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or the processor 51 may also be any conventional processor, etc.
[0341] This application also provides a computer storage medium, please continue to refer to Figure 5 , Figure 5 It is a structural diagram of an embodiment of a computer storage medium provided in the present application. The computer storage medium 600 stores program data 61. When the program data 61 is executed by the processor, it is used to implement the vehicle path planning method of the above embodiment.
[0342] When the embodiments of the present application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.
[0343] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Equivalent structures or equivalent process changes made by utilizing the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A vehicle path planning method, characterized in that: The method comprises: During the driving process of the vehicle, obtain the obstacle vehicle corresponding to the current vehicle; A Monte Carlo search tree is constructed according to the first state of the current vehicle and the second state of the obstacle vehicle; the first state includes the current position, current speed and current vehicle head direction of the current vehicle at the current moment; the second state includes the current position, current speed and current vehicle head direction of the obstacle vehicle at the current moment; wherein the first state is the root node of the Monte Carlo search tree; Determine a plurality of decision branches according to the UCT value corresponding to each node in the Monte Carlo search tree, and determine the best decision branch from the plurality of decision branches, wherein the node is used to simulate the running state of the current vehicle at a future moment; The current vehicle is controlled to travel according to the optimal decision branch.
2. The method according to claim 1, characterized in that The step of constructing a Monte Carlo search tree according to the first state of the current vehicle and the second state of the obstacle vehicle comprises: The root node is used as a parent node to generate at least one child node; the child node is used to simulate the first target state of the current vehicle after a first preset time; and, according to the second state, update the second target state of the obstacle vehicle after the first preset time; wherein the current vehicle in the first target state does not collide with the obstacle vehicle in the second target state; Taking the child node as a parent node, continuing to execute the step of generating at least one child node until the number of layers of the Monte Carlo search tree is greater than or equal to a preset number of layers; When the number of layers of the Monte Carlo search tree is greater than or equal to the preset number of layers, the score corresponding to each node is calculated, and the UCT value corresponding to each node is determined using the score.
3. The method according to claim 2, characterized in that The first state includes a current acceleration of the current vehicle; The first target state includes a target acceleration, a first target heading angle, and a first target curvature of the current vehicle, and the score includes a driving comfort level; The calculating the score corresponding to each node includes: Using the current acceleration and the target acceleration, obtaining the acceleration comfort corresponding to the current node; Obtaining a heading angle comfort degree corresponding to the current node by using the first target heading angle and a second target heading angle of a parent node corresponding to the current node; Using the first target curvature and the second target curvature of the parent node, obtaining the curvature comfort corresponding to the current node; The weighted sum of the acceleration comfort, the heading angle comfort, and the curvature comfort is calculated to obtain the driving comfort of the current node.
4. The method according to claim 2, characterized in that: The first target state includes a target speed and a first target position; The scores include traffic efficiency; The calculating the score corresponding to each node includes: Using the target speed and the maximum speed limit of the road where the vehicle is currently located, the speed efficiency corresponding to the current node is obtained; Obtain all associated nodes between the root node and the current node, and determine the target path length using the position information corresponding to all the associated nodes; The traffic efficiency of the position corresponding to the current node is obtained by using the target longitudinal distance and the target path length; the target longitudinal distance is the longitudinal distance between the first target position and the second target position of the parent node corresponding to the current node; A weighted sum of the speed traffic efficiency and the position traffic efficiency is calculated to obtain the traffic efficiency of the current node.
5. The method according to claim 2, characterized in that: The first target state includes a target acceleration; Said scores include historical scores; The calculating the score corresponding to each node includes: Obtaining a historical acceleration of the current vehicle at a historical moment, and determining a preset acceleration range according to the historical acceleration; In response to the target acceleration being within the preset acceleration interval, a first preset value is added to the historical score.
6. The method according to claim 2, characterized in that The first target state includes a first target position; The score includes a lane change penalty; The calculating the score corresponding to each node includes: Determining whether the current vehicle changes lanes according to a second target position of the target node corresponding to the parent node and the first target position; In response to the current vehicle having a lane change behavior, a second preset value is added to the lane change penalty item.
7. The method according to claim 2, characterized in that Determining a number of decision branches based on the UCT value corresponding to each node in the Monte Carlo search tree includes: Using the root node, constructing the decision branch; Take the root node as the parent node, and add the target child node with the largest UCT value from the child nodes of the next layer to the decision branch; The target child node is taken as the parent node, and the target child node with the largest UCT value is repeatedly added to the decision branch from the next layer of child nodes until the layer number of the target child node is greater than or equal to the preset layer number.
8. The method according to claim 1, characterized in that: The step of obtaining an obstacle vehicle corresponding to the current vehicle during the vehicle driving process includes: Acquire candidate vehicles around the current vehicle; In response to the candidate vehicle colliding with the current vehicle within a second preset time period, the candidate vehicle is used as the obstacle vehicle.
9. A terminal device, characterized in that: The terminal device includes a processor and a memory connected to the processor, wherein: The memory stores program instructions; The processor is configured to execute program instructions stored in the memory to implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The storage medium stores program instructions, and when the program instructions are executed, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Multi-vehicle cooperative lane changing control method, device and equipment and medium
CN120472673A