Planning trajectory generation method and device, electronic equipment and storage medium

By using POMDP models and potential collision risk classification technology in intelligent driving systems, the problem of failure to fully consider the uncertain behavior of traffic participants in the existing technology is solved, and a safer and more efficient planning trajectory generation is achieved.

CN119975408AActive Publication Date: 2025-05-13SAIC GM WULING AUTOMOBILE CO LTD

Patent Information

Application Number
CN202510113155.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The uncertain behavior of other traffic participants is not fully considered in the prior art, resulting in the incomplete final planning trajectory.

Method used

By modeling the decision planning process under the uncertainty of the bypass vehicle behavior into partially observable Markov decision-making process (POMDP), and classifying the predicted trajectory of the bypass vehicle through potential collision risks, reducing the state space to improve decision-making planning efficiency and stability.

Benefits of technology

Effectively predict the uncertain behavior of side cars, ensure the driving safety of the bicycle, and improve the perfection of the planning trajectory and real-time efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119975408A_ABST
    Figure CN119975408A_ABST
Patent Text Reader

Abstract

The invention provides a planned trajectory generation method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining a state space and an action space of a vehicle; according to the potential collision risk and the action space, classifying the predicted trajectories of the side vehicle to obtain a risk trajectory set and a safety trajectory set; and through a partially observable Markov decision process, modeling is performed on a state space, an action space, a risk trajectory set and a safety trajectory set to obtain a reference trajectory. The decision planning process under the uncertainty of the side vehicle behavior is modeled as POMDP, so that the uncertain behavior of the side vehicle is fully predicted, and the driving safety of the vehicle is ensured. Besides, in consideration of too high dimensionality of a multi-traffic-participant prediction trajectory cluster in a complex scene and easy falling into dimensionality disaster of the POMDP model, the prediction trajectory of the side vehicle is classified into a risk trajectory set and a safe trajectory set through a potential collision risk, so that state space dimensionality reduction is realized, and decision planning efficiency and decision stability are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent driving technology, and specifically to a planning trajectory generation method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of automobile technology, more and more cars are equipped with autonomous driving technology. The decision-making and planning module is one of the key parts of autonomous driving. The decision-making and planning module is the process of making some purposeful decisions for a certain goal, usually referring to reaching the destination from the departure point while avoiding obstacles, and constantly optimizing the planned trajectory and behavior to ensure the safety and comfort of passengers.

[0003] In some related technologies, trajectory clusters are directly sampled from the planning space, and then the final planned trajectory is evaluated through indicators such as safety and comfort. However, the behavioral interaction with other traffic participants is not considered, which may lead to the final planned trajectory being incomplete and the trajectory needs to be adjusted during driving.

[0004] In addition, in some related technologies, decision-making methods based on optimization or game theory, although modeling the dynamics of behavioral interactions, only consider the deterministic prediction results of the future motion status of surrounding traffic participants, and lack a response mechanism for uncertain behaviors.

[0005] It should be pointed out that the information disclosed in the background technology section of this application is only intended to deepen the understanding of the general background technology of this application, and should not be regarded as an admission or suggestion in any form that the information constitutes prior art already known to those skilled in the art. Summary of the invention

[0006] In view of this, the present application provides a planning trajectory generation method, device, electronic device and storage medium, so as to solve the problem in the prior art that the final planning trajectory is not perfect due to insufficient consideration of the uncertain behavior of other traffic participants.

[0007] In a first aspect, an embodiment of the present application provides a method for generating a planning trajectory, comprising: Acquire a state space and an action space of the vehicle, wherein the state space includes position information and motion state information of the vehicle and the neighboring vehicle, and the action space includes the driving action of the vehicle; According to the potential collision risk and the action space, the predicted trajectories of the adjacent vehicles are classified to obtain a risk trajectory set and a safe trajectory set; The state space, the action space, the risk trajectory set and the safety trajectory set are modeled through a partially observable Markov decision process to obtain a reference trajectory.

[0008] In the embodiment of the present application, the decision-making and planning process under the uncertainty of the behavior of the adjacent vehicle is modeled as a partially observable Markov decision process (POMDP), so as to fully predict the uncertain behavior of the adjacent vehicle and ensure the driving safety of the vehicle. In addition, considering that the dimension of the predicted trajectory cluster of multiple traffic participants in complex scenes is too high, the POMDP decision-making and planning model is prone to dimensionality disaster, so the predicted trajectory of the adjacent vehicle is classified into a risk trajectory set and a safe trajectory set according to the potential collision risk, thereby achieving state space dimensionality reduction and effectively improving the decision-making and planning efficiency and decision stability.

[0009] In a possible implementation, the action space of the vehicle includes a lateral behavior library and a longitudinal behavior library, and obtaining the action space of the vehicle includes: According to the global plan and lane-level map, a set of feasible lanes is obtained; Extracting lane center lines from lanes in the set of drivable lanes to obtain a reference path set, wherein the reference path set is the lateral behavior library; An acceleration set is obtained, where the acceleration set is the longitudinal behavior library.

[0010] In the embodiment of the present application, according to the global planning and lane-level map, a set of drivable lanes is obtained, and the lane centerline is extracted from the lanes in the drivable lane set to obtain a reference path set, which is a lateral behavior library, and the acceleration set is used as a longitudinal behavior library according to the typical longitudinal motion mode. It can be understood that the lateral behavior library in the present application uses a reference path set rather than a lane set. Compared with the lane, the lane centerline as a path is easier to calculate, which helps to reduce the burden of the algorithm.

[0011] In a possible implementation, before classifying the predicted trajectory of the adjacent vehicle according to the potential collision risk and the action space, the method further includes: Probability calculation is performed on the predicted trajectory of the adjacent vehicle to obtain the predicted trajectory probability distribution; According to the predicted trajectory probability distribution, predicted trajectories having a probability less than or equal to a preset probability threshold are deleted.

[0012] In the embodiment of the present application, before processing the predicted trajectory of the adjacent vehicle, the predicted trajectory of the adjacent vehicle is firstly calculated for probability to obtain the predicted trajectory probability distribution, and according to the predicted trajectory probability distribution, the predicted trajectory with a probability less than or equal to a preset probability threshold is deleted. It can be understood that by performing probability calculation on the predicted trajectory in advance, the trajectory that is almost impossible to occur can be eliminated in advance, which will not affect the decision safety and will greatly optimize the real-time performance.

[0013] In a possible implementation, the predicted trajectories of adjacent vehicles are classified according to the potential collision risk and the action space to obtain a risk trajectory set and a safety trajectory set, including: Determining a driving lane for the vehicle according to the action space; Determine whether the predicted trajectory of the adjacent vehicle intersects with the lane of the own vehicle; If the predicted trajectory of any of the adjacent vehicles intersects with the lane of the vehicle, the predicted trajectory of any of the adjacent vehicles is determined to be a risk trajectory; If the predicted trajectory of any of the adjacent vehicles does not intersect with the lane of the vehicle, the predicted trajectory of any of the adjacent vehicles is determined to be a safe trajectory.

[0014] In the embodiment of the present application, by judging whether the predicted trajectory of the adjacent vehicle intersects with the lane of the vehicle, it is determined whether the predicted trajectory of the adjacent vehicle is a risky trajectory or a safe trajectory. It can be understood that by converting the probability distribution of the predicted trajectory into the probability distribution of whether there is a potential collision risk with the vehicle, the size of the confidence state space is directly reduced to 2 states, which reduces the solution time of the POMDP and improves the real-time efficiency.

[0015] In a possible implementation, the reference trajectory is a trajectory generated with the lane centerline as a reference, and after obtaining the reference trajectory, the method further includes: The reference trajectory is adjusted laterally to obtain a planned trajectory.

[0016] In the embodiment of the present application, the reference trajectory is a trajectory generated with reference to the center line of the lane. However, the vehicle obviously cannot travel along the center line of the lane when driving, so it is necessary to adjust the reference trajectory laterally to obtain a planned trajectory that is more in line with actual driving habits.

[0017] In a possible implementation manner, the step of laterally adjusting the reference trajectory to obtain a planned trajectory includes: Constructing a Frenet coordinate system using the tangent vector and the normal vector of the reference trajectory; In the Frenet coordinate system, a planned trajectory is generated at a position at a preset distance from the reference trajectory.

[0018] In the embodiment of the present application, the planned trajectory is generated at a position at a preset distance from the reference trajectory through the Frenet coordinate system. It can be understood that the Frenet coordinate system can more intuitively represent the position of the vehicle on a curved road, which is more conducive to generating a planned trajectory that meets expectations.

[0019] In a possible implementation, the state space, the action space, the risk trajectory set, and the safety trajectory set are modeled through a partially observable Markov decision process to obtain a reference trajectory, including: The state space, the action space, the risk trajectory set and the safety trajectory set are modeled through a partially observable Markov decision process to obtain a feasible trajectory set; A reference trajectory is obtained from the feasible trajectory set through a reward function.

[0020] In the embodiment of the present application, after POMDP modeling, a feasible trajectory set is obtained, and a reference trajectory is obtained from the feasible trajectory set through a reward function. It can be understood that the feasible trajectory set includes multiple trajectories, and the planning trajectory generation device needs to select a safe, reasonable, passenger comfort-enhancing and highly efficient trajectory from the multiple trajectories as a reference trajectory.

[0021] In a second aspect, an embodiment of the present application provides a planning trajectory generation device, including: An information acquisition module, used to acquire a state space and an action space of the ego vehicle, wherein the state space includes position information and motion state information of the ego vehicle and adjacent vehicles, and the action space includes driving actions of the ego vehicle; A predicted trajectory classification module is used to classify the predicted trajectories of adjacent vehicles according to the potential collision risk and the action space to obtain a risk trajectory set and a safe trajectory set; The modeling module is used to model the state space, the action space, the risk trajectory set and the safety trajectory set through a partially observable Markov decision process to obtain a reference trajectory.

[0022] In a third aspect, an embodiment of the present application provides an electronic device, including: processor; Memory; And a computer program, wherein the computer program is stored in the memory, and the computer program includes instructions, and when the instructions are executed by the processor, the electronic device executes any one of the methods described in the first aspect.

[0023] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, characterized in that the computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the methods described in the first aspect.

[0024] It is understandable that the planning trajectory generation device provided in the second aspect, the electronic device provided in the third aspect, and the computer-readable storage medium provided in the fourth aspect are used to execute the method provided in the present application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, which will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0026] Figure 1 A schematic diagram of a flow chart of a planning trajectory generation method provided in an embodiment of the present application; Figure 2 A lane schematic diagram of a lateral action library provided in an embodiment of the present application; Figure 3 A schematic diagram of a potential collision risk assessment provided in an embodiment of the present application; Figure 4 A schematic diagram of probability calculation of a representative trajectory provided in an embodiment of the present application; Figure 5 A schematic diagram of a lane-changing trajectory of a vehicle provided in an embodiment of the present application; Figure 6 A schematic diagram of the architecture of another planning trajectory generation method provided in an embodiment of the present application; Figure 7 A schematic diagram of the structure of a planning trajectory generation device provided in an embodiment of the present application; Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] In order to better understand the technical solution of the present application, the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0028] It should be clear that the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.

[0029] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0030] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0031] With the rapid development of automobile technology, more and more cars are equipped with autonomous driving technology. The decision-making and planning module is one of the key parts of autonomous driving. The decision-making and planning module is the process of making some purposeful decisions for a certain goal, usually referring to reaching the destination from the departure point while avoiding obstacles, and constantly optimizing the planned trajectory and behavior to ensure the safety and comfort of passengers.

[0032] In some related technologies, trajectory clusters are directly sampled from the planning space, and then the final planned trajectory is evaluated through indicators such as safety and comfort. However, the behavioral interaction with other traffic participants is not considered, which may lead to the final planned trajectory being incomplete and the trajectory needs to be adjusted during driving.

[0033] In addition, in some related technologies, decision-making methods based on optimization or game theory, although modeling the dynamics of behavioral interactions, only consider the deterministic prediction results of the future motion status of surrounding traffic participants, and lack a response mechanism for uncertain behaviors.

[0034] In response to the above problems, an embodiment of the present application provides a planning trajectory generation method, which models the decision-making planning process under the uncertainty of the behavior of the adjacent vehicle as a partially observable Markov decision process (POMDP), so as to fully predict the uncertain behavior of the adjacent vehicle and ensure the driving safety of the vehicle. In addition, considering that the dimension of the predicted trajectory cluster of multiple traffic participants in complex scenes is too high, the POMDP decision-making planning model is prone to fall into the dimensionality curse, so the predicted trajectory of the adjacent vehicle is classified into a risk trajectory set and a safe trajectory set according to the potential collision risk, thereby achieving state space dimensionality reduction and effectively improving the decision-making planning efficiency and decision stability. This is described in detail in conjunction with the accompanying drawings and specific embodiments below.

[0035] See also Figure 1 , is a flow chart of a planning trajectory generation method provided in an embodiment of the present application. Figure 1 As shown, it mainly includes the following steps.

[0036] Step S101: Obtain the state space and the action space of the vehicle.

[0037] Among them, the state space includes the position information and motion state information of the vehicle and the adjacent vehicle. Specifically, the state space includes the position (x, y), driving direction θ, speed v, driving distance s along the target reference path, and driving intention I of the adjacent vehicle. Driving intention I is used to characterize the future trajectory of the adjacent vehicle, which is the part of the system state that cannot be directly observed. In the embodiment of the present application, it is assumed that the driving intention I of the adjacent vehicle is a certain trajectory in the predicted trajectory cluster output by the prediction model, which is uncertain at the time of decision-making, and the influence of this uncertainty factor on the collision risk assessment is investigated through Monte Carlo sampling. The POMDP decision planning model needs to find the optimal strategy that satisfies the maximum discounted cumulative return under the probability distribution of driving intention. Other road information and fixed buildings are used as static reference objects and are not added to the state space.

[0038] The state definitions in the embodiment of the present application are as follows:

[0039]

[0040]

[0041] Among them, s ego Indicates the state of the vehicle. It represents the status of the neighboring vehicles around the vehicle, and n is the number of neighboring vehicles.

[0042] In the embodiment of the present application, the action space includes the driving action of the vehicle. Specifically, the action space includes two parts: a longitudinal action library and a lateral action library.

[0043] Among them, the longitudinal behavior storage location acceleration set, in this embodiment, the longitudinal behavior library includes 4 discrete accelerations, namely:

[0044] Among them, A lon is the longitudinal motion of the vehicle. It can be understood that when A lon for When the acceleration is 0m / s 2 , at this time the vehicle is moving at a constant speed; when A lon for When the acceleration is 1.2 m / s 2 , at this time the vehicle accelerates; when A lon for When the acceleration is -1.8 m / s 2 , at this time the vehicle slows down; when A lon for When the acceleration is -3.6 m / s 2 , the vehicle is forced to brake at this time.

[0045] In some possible implementations, when the situation ahead is too urgent, even if the acceleration corresponding to the strong braking in the above longitudinal action is adopted, the front collision cannot be avoided, then the emergency braking mode is entered and the following longitudinal action is adopted for POMDP replanning:

[0046] The emergency braking mode not only ensures the safety of emergency braking, but also avoids the continuous expansion of the action space, which in turn affects the efficiency of POMDP solution.

[0047] In some related technologies, only longitudinal speed planning is performed, while the lateral trajectory is planned using other methods outside the POMDP model. In some other related technologies, although both longitudinal speed and lateral trajectory planning are modeled in the POMDP model, the lateral trajectory is fixedly defined as driving in a predetermined mode, including left turns, lane keeping, right lane changes, going straight at intersections, turning left at intersections, and right turn lights at intersections. This method of pre-defining driving actions based on traffic scenarios may not be able to traverse all required traffic scenarios, and these redundant actions will increase or decrease the dimension of the POMDP action space, directly leading to an exponential increase in the time complexity of POMDP solution.

[0048] To address this problem, in an embodiment of the present application, the lateral behavior library is dynamically adjusted to a set of currently drivable reference paths, and its action space is more in line with real-time traffic scenarios.

[0049] Specifically, in the lane-level map, guided by the global planning path, the current drivable lane is determined and its lane centerline is read, and then a reference path is generated through a piecewise cubic B-spline curve.

[0050] First, determine the preview distance within the decision cycle based on the current speed and maximum acceleration. The formula is as follows:

[0051] In the formula, is the current speed of the vehicle, is the maximum acceleration allowed, and T is the duration of the decision horizon.

[0052] Then intercept the path points within the preview distance of the current lane and the target lane as the control points of the spline curve Each small segment of cubic B-spline curve consists of 4 control points P i , P i+1 , P i+2 and P i+3 Generated, piecewise cubic B-spline curve can be described as:

[0053] Finally, the generated horizontal action library , where path i (i=0, 1, ..., n) is the drivable lane of the vehicle. For ease of understanding, the embodiment of the present application takes two drivable lanes as an example to explain the drivable lanes in detail.

[0054] See also Figure 2 , is a lane diagram of a lateral action library provided in an embodiment of the present application. Figure 2 As shown, the vehicle is traveling straight in one lane of a dual-lane road. In this embodiment, the first path path in the lateral action library 0 To go straight along the current lane, the second path in the lateral action library is 1 To change lane to the left.

[0055] The action space A of the POMDP model can be obtained by combining the elements in the vertical action library and the horizontal action library in pairs. The specific formula is as follows:

[0056] It should be noted that the longitudinal motion in the formula of the embodiment of the present application includes , , and However, in practical applications, the longitudinal action may also include other actions, and the embodiments of the present application do not impose specific limitations on this.

[0057] In addition, in most cases, the current drivable paths are within 3, namely the current lane and the left and right adjacent lanes. Therefore, the dimension of the lateral action library changes dynamically in the range of 0-3 in most cases, and the dimension of the corresponding action space A changes dynamically in the range of 0-12. In order to avoid the vehicle changing lanes back and forth within the decision cycle, the embodiment of the present application limits the operation of changing the lateral path to only one time in each decision cycle, thereby saving solution time and improving work efficiency.

[0058] Step S102: Classify the predicted trajectories of the adjacent vehicles according to the potential collision risk and the action space to obtain a risk trajectory set and a safe trajectory set.

[0059] Specifically, the driving lane of the vehicle is determined according to the action space, and it is judged whether the predicted trajectory of the adjacent vehicle intersects with the driving lane of the vehicle. If the predicted trajectory of any adjacent vehicle intersects with the driving lane of the vehicle, the predicted trajectory is determined to be a risky trajectory; if the predicted trajectory of any adjacent vehicle does not intersect with the driving trajectory of the vehicle, the predicted trajectory is determined to be a safe trajectory.

[0060] To facilitate understanding, an embodiment of the present application provides a schematic diagram for evaluating potential collision risks.

[0061] See also Figure 3 , is a schematic diagram of a potential collision risk assessment provided by an embodiment of the present application. Figure 3 As shown in the figure, the red vehicle is the adjacent vehicle, and the two adjacent vehicles are adjacent vehicle A and adjacent vehicle B. The blue vehicle is the ego vehicle. Taking a two-lane road as an example, the ego vehicle has two driving decision candidates: (a) lane keeping and (b) lane changing. In each decision, the center line of the target lane is used as the local planning reference path, and then the prediction trajectory cluster is judged to determine whether it intersects with the lane area occupied by the reference path to determine whether there is a potential collision risk. Figure 3 As shown, the reference path is a blue dotted line, the center lines of other lanes are black dotted lines, the blue trajectory in the predicted trajectory cluster of the adjacent vehicle is a safe trajectory, and the red trajectory is a risky trajectory.

[0062] like Figure 3 As shown in (a) in the figure, the vehicle is currently driving in the right lane. When the vehicle maintains the current lane, the trajectories in the predicted trajectory clusters of the adjacent vehicles A and B that intersect with the right lane are all risk trajectories, and the trajectories in the predicted trajectory clusters of the adjacent vehicles A and B that do not intersect with the right lane are all safe trajectories. Similarly, if Figure 3 As shown in (b) in the figure, the vehicle is currently traveling in the right lane, but the vehicle's driving decision is to change lanes, so the reference path corresponding to the vehicle is the center line of the left lane. Since all trajectories in the predicted trajectory clusters of the adjacent vehicles A and B intersect with the left lane, all predicted trajectories of the adjacent vehicles A and B are risk trajectories.

[0063] It is understandable that the above classification of the potential collision risk of the predicted trajectory of the adjacent vehicle is static and has nothing to do with the future speed of the vehicle itself. This is because the speed of the vehicle itself is obtained by solving the POMDP decision model and is unknown at this time. However, in the embodiment of the present application, if a certain predicted trajectory is determined to be a safe trajectory, it means that the vehicle itself will not collide with the adjacent vehicle at any speed.

[0064] After the predicted trajectories of the adjacent vehicles are divided into two categories according to whether they cause potential collision risks to a candidate strategy of the vehicle: risky trajectories and safe trajectory , and calculate the sum of the probabilities of the corresponding categories. Then, and The trajectory with the highest probability is selected as the representative trajectory to participate in the solution of collision safety. Specifically, the calculation method is as follows:

[0065]

[0066] Wherein, P is the probability, and track is the representative track of the category, which is used for sampling and as the future track for constructing the state transition model of the surrounding vehicle i. Figure 3 Taking the case where the vehicle is driving straight in the right lane and there is a neighboring vehicle B in the left lane as an example, the probability calculation of the representative trajectory is described in detail.

[0067] See also Figure 4 , is a schematic diagram of probability calculation of a representative trajectory provided in an embodiment of the present application. Figure 4 As shown in the figure, the predicted trajectory cluster of the adjacent vehicle B includes 495 trajectories, that is, the total number of blue safe trajectories and red risk trajectories is 495, among which the maximum probability safe trajectory (i.e. safe trajectory Representative trajectory of ) is a bold blue line, and the maximum probability risk trajectory (i.e. risk trajectory Representative trajectory of ) is the bold red line, which can be obtained by calculation The predicted probability is 0.137, The prediction probability of is 0.863. The threat level of the adjacent vehicle to the vehicle can be inferred through the prediction probabilities corresponding to the safe trajectory and the risky trajectory.

[0068] It is understandable that there are usually dozens or even hundreds of predicted trajectory clusters of adjacent vehicles. If they are directly input into the POMDP model and used as the future states of adjacent vehicles for collision detection, the computational load of the POMDP model is very large. The predicted trajectory classification method provided in the embodiment of the present application greatly reduces the complexity of the solution, and converts the confidence space of dozens or hundreds of predicted trajectory clusters into a confidence space of only two states of whether there is a collision risk, thereby improving work efficiency.

[0069] As mentioned above, there are hundreds or thousands of predicted trajectories of neighboring vehicles, and many of them have very low probability. If these extremely low probability trajectories are treated equally and allowed to participate in subsequent calculations, it is meaningless and will seriously deteriorate the real-time performance of the decision-making system.

[0070] Therefore, in a possible implementation, before classifying the predicted trajectories of the adjacent vehicles, the method further includes performing probability calculation on the predicted trajectories of the adjacent vehicles to obtain a predicted trajectory probability distribution; based on the predicted trajectory probability distribution, the predicted trajectories with a probability less than or equal to a preset probability threshold are deleted. The preset probability threshold is a fixed threshold set based on debugging experience, and those skilled in the art can set it to any value based on actual needs, and the present application embodiment does not impose specific restrictions on this.

[0071] It can be understood that the trajectory with extremely low probability is the trajectory that is almost impossible to occur. For example, the historical trajectory of the adjacent car always keeps going straight in the current lane. Therefore, in the prediction trajectory cluster output by the prediction module, the probability of the trajectory cluster corresponding to the left turn is very small, indicating that the left turn is an extremely low probability event. If these trajectories are eliminated, it will not affect the decision safety and will greatly optimize the real-time performance.

[0072] Step S103: Modeling the state space, action space, risk trajectory set and safety trajectory set through a partially observable Markov decision process to obtain a reference trajectory.

[0073] Specifically, in order to obtain the state of the vehicle and the adjacent vehicles, the state transition model is used in the embodiment of the present application to describe the evolution of the state of each participant in the traffic scene over time. The state transition model is a state transition probability. The embodiment of the present application involves the behavioral interaction between the vehicle and the neighboring vehicle, so the embodiment of the present application describes the state transfer model from the perspective of the vehicle and the neighboring vehicle respectively.

[0074] The embodiment of the present application mainly studies the decision-making problem under the uncertainty of the driving intention of the adjacent vehicle. It is assumed that the motion state information such as the position, speed and heading of the self-vehicle and the adjacent vehicle is accurate, and the error magnitude is small enough not to affect the decision result. In the system, only the driving intention of the adjacent vehicle is uncertain. In addition, since the decision-making module is more abstract than the control module, the accuracy requirements for positioning and state evolution are not very high, so the state transition model of the self-vehicle can be completely determined by the vehicle kinematic model, and there is no need to involve a complex calculation of the dynamic model. From the definition of the action space, it can be seen that the decision output of the self-vehicle is first to select a drivable path dynamically generated outside the model, and then perform longitudinal movement on the selected lateral path. Therefore, the state transition model of the vehicle can be expressed as:

[0075] In the formula, is the acceleration corresponding to the longitudinal action taken at a decision step, represents the drivable path i corresponding to the lateral action taken in this decision step, and S is the path along The distance traveled, the initial value is The starting point and the current point of the vehicle are The distance between the projections on the selected path, x, y, θ The current driving distance S is interpolated and △t is the decision step length.

[0076] The state transition model of the side car is different from that of the ego car. The ego car state update is along the selected lateral path. The surrounding vehicles do not involve changes in driving actions, and the predicted trajectory is a series of time-varying Trajectory point, so its state update is obtained directly on the predicted trajectory according to the time interpolation of each decision step. However, the driving intention I represented by the predicted trajectory cannot obtain the true value, and can only be evaluated and decided on its probability distribution. To this end, the embodiment of the present application randomly samples a predicted trajectory cluster from the probability distribution of driving intention I, and uses it as the future trajectory of the adjacent car in a certain decision-making process. Calculating the discounted cumulative return average after multiple Monte Carlo sampling simulations can approximate the optimal decision under the probability distribution of driving intention I. Therefore, in a certain Monte Carlo sampling simulation, driving intention I is unique and deterministic, that is, the predicted trajectory sampled this time. At this point, the state transition model of adjacent car i can be expressed as the predicted trajectory sampled The current decision step time is directly Interpolate to get new trajectory points:

[0077] In the formula, For the safety trajectory described above, is the risk trajectory described above.

[0078] In addition, in the embodiment of the present application, the posterior probability distribution of the current state is estimated by the observation quantity z, thereby obtaining an observation space Z similar to the state space, which is defined as:

[0079]

[0080]

[0081] in, Represents the measurement value of the vehicle, including the vehicle position, heading angle and speed information. It represents the measured value of the adjacent vehicle, including the position, heading angle, speed and driving intention of the adjacent vehicle. The driving intention of the adjacent vehicle I is the probability distribution of the future trajectory of the predicted trajectory cluster after the potential collision risk clustering and dimensionality reduction.

[0082] The observation model O indicates that after taking action a in the current state, it reaches a new state The probability of the measured value at . The sensing system in the embodiment of the present application is based on high-precision combined inertial navigation positioning and high-precision laser radar. Both the vehicle and the adjacent vehicle have centimeter-level positioning accuracy. The sensing error level is much smaller than the impact of the uncertainty of the adjacent vehicle's intention on the decision result. Therefore, the embodiment of the present application does not consider the observation error, and considers that the measured value is equivalent to its corresponding state value, that is, after executing action a, the measured value of the vehicle changes completely with the state transfer model. The state and measurement value of the adjacent vehicle are determined by the future predicted trajectory. Although its driving intention cannot be directly observed, its predicted trajectory probability distribution can be output through the prediction module. Then, within the decision cycle, after fully considering its probability distribution and uncertainty, the best future trajectory is sampled as the basis for updating the state and measurement value of the adjacent vehicle.

[0083] In one possible implementation, in order to obtain a reference trajectory with greater benefits, the state space, action space, risk trajectory set and safety trajectory set are modeled through POMDP. After obtaining the feasible trajectory set, the reference trajectory is obtained in the feasible trajectory set through the reward function.

[0084] Specifically, the reward function is the immediate return after executing action a in confidence state b. Generally, multiple evaluation indicators are used to quantify and compare the pros and cons of action a in state b. That is, by accumulating the immediate rewards in the entire decision cycle, the state action sequence with the highest discounted cumulative reward is selected. , Security , Comfort and driving efficiency The reward function is designed in several aspects, and the overall reward function and the effects of these aspects are:

[0085] The global task constraints are as follows: After the global path is given, the preview distance is determined according to the formula The local reference path can be intercepted from the global reference path, and the end point of the local reference path is used as the target point of the current decision cycle. Under the premise of ensuring driving safety and comfort, the longer the distance traveled along the local reference path, the closer to the target point, the higher the task completion degree under the guidance of the global path, and the greater the benefit. The formula is as follows:

[0086] in, and is the weight coefficient, is the distance from the projection point of the local reference path to the target point of the vehicle, is the lateral vertical distance from the actual position of the vehicle to the local reference path. Unreasonable lane changing actions can be suppressed through global task constraints.

[0087] The specific driving safety constraint is: starting from the current moment, evaluate the potential safety of whether the ego vehicle will collide with the adjacent vehicle in the future decision cycle. According to whether the planned trajectory of the ego vehicle and the predicted trajectory of the adjacent vehicle collide within a certain decision step time interval △t, the driving safety is evaluated. The formula is as follows:

[0088] in, is the weight coefficient, relative to other weights, Generally, a larger constant, such as 1000, is used to penalize actions that cause collisions. is the speed of the vehicle at the time of the collision, and the constant Used to ensure penalty when the vehicle speed is 0.

[0089] The comfort constraint is as follows: Occupant comfort is an important indicator for measuring intelligent driving systems. Current research mainly characterizes it by acceleration and its rate of change. Acceleration includes two dimensions: longitudinal and lateral acceleration: , where v is the longitudinal speed, θ is the vehicle heading angle, is the longitudinal acceleration, is the lateral acceleration. Excessive acceleration change rate j is reflected in the operation level as frequent acceleration and deceleration, which will cause great discomfort. When the acceleration and its change rate exceed the critical value, the comfort deteriorates as the two increase, as follows:

[0090] in, , , and is the corresponding weight coefficient.

[0091] The driving efficiency constraint is as follows: In order to encourage intelligent driving vehicles to travel at the expected speed while ensuring safety and comfort, the speed constraint items are set as follows:

[0092] in, is the maximum speed allowed by traffic regulations, is the weight coefficient.

[0093] It should be pointed out that all the above-mentioned weight coefficients are numerical values ​​set based on experience. Those skilled in the art can set the weight coefficients to any numerical values ​​according to actual needs, and the embodiments of the present application do not impose specific limitations on this.

[0094] After the POMDP decision-making planning model is constructed and the state space dimension is reduced to improve decision-making efficiency and stability, the POMDP model is solved using the DESPOT solver, and finally a reference trajectory along the center line of the road containing speed information is output. The speed is obtained by integrating the longitudinal action sequence, that is, the expected acceleration. The reference path corresponding to the reference trajectory is the drivable path selected by the lateral action sequence, and its terminal end point is only located in the center of the corresponding lane.

[0095] join Figure 5 , is a schematic diagram of a lane-changing trajectory of a vehicle provided in an embodiment of the present application. Figure 5 As shown, the terminal end of the reference trajectory is located at the center of the corresponding lane. The purpose of this is to reduce the dimension of the POMDP action space, but it brings a problem. The vehicle always drives along the center line of the lane, which is obviously not in line with actual driving habits. Therefore, based on the reference trajectory, the embodiment of the present application keeps the speed information unchanged, samples more planning trajectory sets horizontally, and selects the optimal planning trajectory according to the performance index, thereby improving the rationality of the planning trajectory.

[0096] Specifically, the reference trajectory output by POMDP is used as the reference line, and the tangent vector and normal vector of this line are used to construct the Frenet coordinate system. In the Frenet coordinate system, the horizontal coordinate s describes the distance along the reference line, and the vertical coordinate d represents the displacement of the normal vector of the current position along the reference line, that is, the distance from the reference line. The Frenet coordinate system is easier to describe how far the vehicle has traveled along the curved lane and whether it has deviated from the center of the lane than the Cartesian coordinate system. Keeping the length of the reference trajectory unchanged, a set of terminal normal displacements d of different sizes are sampled in the target lane, and combined with the starting state, a set of quintic polynomial curves covering the lateral space of the target lane can be solved. It is worth noting that in order not to destroy the existing collision safety of the reference trajectory, the sampled d must satisfy the following formula, otherwise, the newly sampled planning trajectory set needs to be evaluated for collision risk under the uncertainty of the intentions of surrounding vehicles, so the reference trajectory solved by POMDP previously loses its meaning.

[0097]

[0098] in, is the vehicle width, For safety reasons, some lateral space is left on the vehicle. is the lateral safety distance threshold.

[0099] Different from most of the current methods that use Frenet coordinate sampling for trajectory planning, the reference line here is not a global path, but a reference trajectory that has been solved through POMDP modeling and has met the collision safety, driving efficiency, comfort and other indicators as much as possible. Here, only the reference trajectory is refined and optimized. Therefore, the distance s along the reference line is no longer sampled, and only the end normal displacement d is sampled, so the size of the generated planning trajectory candidate set is very small and the amount of calculation is not large.

[0100] After obtaining the above planning trajectory candidate set, it is necessary to calculate the performance of each candidate planning trajectory according to the cost evaluation function, and then select the optimal trajectory as the planning trajectory output to the vehicle control module. Specifically:

[0101] In the formula, is the loss function term, is the corresponding weight coefficient, is the most likely trajectory index, N is the total number of candidate trajectories. Since the collision safety of the reference trajectory has been guaranteed by POMDP modeling, combined with the formula to limit the sampling range of the lateral distance d, there is no need to perform collision risk assessment here.

[0102] Loss Term J d Indicates the degree of deviation from the reference trajectory, the purpose is to make the vehicle tend to drive along the center of the road in most cases. d The specific calculation method is:

[0103] Where, d i is the sampling deviation distance of candidate trajectory i, d max is the maximum deviation distance.

[0104] Trajectory curvature smoothing cost J k This is to improve ride comfort and avoid violent lateral movement shocks. k The specific calculation method is:

[0105] In the formula, τ i refers to candidate trajectory i, k is the curvature of the trajectory point.

[0106] Loss Term J c This is to punish the discontinuity of the two planned trajectories and avoid unstable phenomena such as vehicle oscillation and overshoot. c The specific calculation method is:

[0107] In the formula, It is the terminal deviation distance corresponding to the optimal trajectory planned last time.

[0108] Corresponding to the above embodiment, the present application also provides another planning trajectory generation method.

[0109] See also Figure 6 , is a schematic diagram of the architecture of another planning trajectory generation method provided in an embodiment of the present application. Figure 6 As shown, the state space is obtained through the positioning system, the feasible lane set is obtained through global planning and lane-level map, and the reference path set is obtained according to the feasible lane set. The reference path set can constitute the lateral behavior library, the acceleration set is the longitudinal behavior library, and the lateral behavior library and the longitudinal behavior library constitute the action space.

[0110] The planning trajectory generation device obtains the predicted trajectory of the adjacent vehicle, calculates the probability distribution of the predicted trajectory, prunes the probability distribution of the predicted trajectory (i.e., deletes the predicted trajectory with a probability less than or equal to the preset probability threshold), and classifies the predicted trajectory according to the reference path set, and divides the reference path into risky trajectory and safe trajectory. Specifically, the risky trajectory and the safe trajectory are summed up respectively, and the probability of each trajectory is calculated, and the maximum value of the risky trajectory and the safe trajectory is taken as the representative trajectory. The state transition of the adjacent vehicle is obtained by probabilistically sampling two representative trajectories. The safety distance is dynamically adjusted, and the collision risk is evaluated based on the distribution ST graph, and the action corresponding to the maximum benefit is obtained through the reward function.

[0111] The state space, action space, state transfer of the adjacent vehicle and reward function are modeled by POMDP, and the lateral and longitudinal behavior sequences are solved by DESPOT to generate a reference trajectory. The reference trajectory is then sampled and optimized laterally to obtain the planned trajectory.

[0112] Corresponding to the above embodiments, the present application also provides a planning trajectory generation device.

[0113] See also Figure 7 , is a schematic diagram of the structure of a planning trajectory generation device provided in an embodiment of the present application. Figure 7 As shown, the planning trajectory generation device includes: an information acquisition module 701, a prediction trajectory classification module 702, and a modeling module 703. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the electronic device shown in the figure does not constitute a limitation on the embodiments of the present application. It can be a bus structure or a star structure, and can also include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0114] The information acquisition module 701 is used to acquire the state space and the action space of the vehicle, wherein the state space includes the position information and motion state information of the vehicle and the neighboring vehicles, and the action space includes the driving action of the vehicle; A predicted trajectory classification module 702 is used to classify the predicted trajectories of adjacent vehicles according to the potential collision risk and the action space to obtain a risk trajectory set and a safe trajectory set; The modeling module 703 is used to model the state space, the action space, the risk trajectory set and the safety trajectory set through a partially observable Markov decision process to obtain a reference trajectory.

[0115] Corresponding to the above embodiments, the present application also provides an electronic device.

[0116] See also Figure 8 , is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 8 As shown, the electronic device 800 may include: a processor 801, a memory 802 and a communication unit 803. These components communicate via one or more buses. Those skilled in the art will appreciate that the structure of the electronic device shown in the figure does not limit the embodiments of the present application. It may be a bus structure or a star structure, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0117] The communication unit 803 is used to establish a communication channel so that the electronic device can communicate with other devices, receive user data sent by other devices or send user data to other devices.

[0118] The processor 801 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. It runs or executes software programs, instructions, and / or modules stored in the memory 802, and calls data stored in the memory to perform various functions of the electronic device and / or process data. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions. For example, the processor 801 can only include a central processing unit (CPU). In the embodiment of the present application, the CPU can be a single computing core or multiple computing cores.

[0119] The memory 802 is used to store the execution instructions of the processor 801. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0120] When the execution instruction in the memory 802 is executed by the processor 801, the electronic device 800 can execute Figure 1 Some or all of the steps in the illustrated embodiments.

[0121] In a specific implementation, the embodiment of the present application further provides a computer storage medium, wherein the computer storage medium may store a program, and when the program is executed, the program may include some or all of the steps in each embodiment of the simulation scene generation method provided in the embodiment of the present application. The storage medium may be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0122] In a specific implementation, an embodiment of the present application also provides a computer program product, wherein the computer program product includes executable instructions, and when the executable instructions are executed on a computer, the computer executes part or all of the steps in each embodiment of the simulation scene generation method provided in the embodiment of the present application.

[0123] In the embodiments of the present application, "at least one" refers to one or more, and "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0124] Those of ordinary skill in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented in a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0125] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0126] In several embodiments provided in the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0127] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the device embodiment and the terminal embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.

Claims

1. A planning trajectory generation method, characterized in that: include: Acquire a state space and an action space of the vehicle, wherein the state space includes position information and motion state information of the vehicle and the neighboring vehicle, and the action space includes the driving action of the vehicle; According to the potential collision risk and the action space, the predicted trajectories of the adjacent vehicles are classified to obtain a risk trajectory set and a safe trajectory set; The state space, the action space, the risk trajectory set and the safety trajectory set are modeled through a partially observable Markov decision process to obtain a reference trajectory.

2. The method according to claim 1, characterized in that The action space of the vehicle includes a lateral behavior library and a longitudinal behavior library. The step of obtaining the action space of the vehicle includes: According to the global plan and lane-level map, a set of feasible lanes is obtained; Extracting lane center lines from lanes in the set of drivable lanes to obtain a reference path set, wherein the reference path set is the lateral behavior library; An acceleration set is obtained, where the acceleration set is the longitudinal behavior library.

3. The method according to claim 1, characterized in that Before classifying the predicted trajectory of the adjacent vehicle according to the potential collision risk and the action space, the method further includes: Probability calculation is performed on the predicted trajectory of the adjacent vehicle to obtain the predicted trajectory probability distribution; According to the predicted trajectory probability distribution, predicted trajectories having a probability less than or equal to a preset probability threshold are deleted.

4. The method according to claim 1, characterized in that The method of classifying the predicted trajectories of adjacent vehicles according to the potential collision risk and the action space to obtain a risk trajectory set and a safety trajectory set includes: Determining a driving lane for the vehicle according to the action space; Determine whether the predicted trajectory of the adjacent vehicle intersects with the lane of the own vehicle; If the predicted trajectory of any of the adjacent vehicles intersects with the lane of the vehicle, the predicted trajectory of any of the adjacent vehicles is determined to be a risk trajectory; If the predicted trajectory of any of the adjacent vehicles does not intersect with the lane of the vehicle, the predicted trajectory of any of the adjacent vehicles is determined to be a safe trajectory.

5. The method according to claim 1, characterized in that The reference trajectory is a trajectory generated with reference to the lane centerline. After obtaining the reference trajectory, the method further includes: The reference trajectory is adjusted laterally to obtain a planned trajectory.

6. The method according to claim 5, characterized in that The step of laterally adjusting the reference trajectory to obtain a planned trajectory includes: Constructing a Frenet coordinate system using the tangent vector and the normal vector of the reference trajectory; In the Frenet coordinate system, a planned trajectory is generated at a position at a preset distance from the reference trajectory.

7. The method according to claim 1, characterized in that The method of modeling the state space, the action space, the risk trajectory set, and the safety trajectory set through a partially observable Markov decision process to obtain a reference trajectory includes: The state space, the action space, the risk trajectory set and the safety trajectory set are modeled through a partially observable Markov decision process to obtain a feasible trajectory set; A reference trajectory is obtained from the feasible trajectory set through a reward function.

8. A planning trajectory generation device, characterized in that: include: An information acquisition module, used to acquire a state space and an action space of the ego vehicle, wherein the state space includes position information and motion state information of the ego vehicle and adjacent vehicles, and the action space includes driving actions of the ego vehicle; A predicted trajectory classification module is used to classify the predicted trajectories of adjacent vehicles according to the potential collision risk and the action space to obtain a risk trajectory set and a safe trajectory set; The modeling module is used to model the state space, the action space, the risk trajectory set and the safety trajectory set through a partially observable Markov decision process to obtain a reference trajectory.

9. An electronic device, characterized in that: include: processor; Memory; And a computer program, wherein the computer program is stored in the memory, and the computer program includes instructions, and when the instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Traffic vehicle intention recognition method based on driving behavior generation mechanism

    CN113911129A

  • Surrounding vehicle trajectory prediction method applied to autonomous vehicle

    CN114872727A

  • Automatic driving automobile lane changing decision control method considering uncertainty

    CN115257746A

  • Intelligent vehicle decision-making method considering uncertainty at intersection

    CN116373905A

  • Unmanned vehicle navigation decision planning system and method based on partial Markov decision process

    CN117870689A

Cited By

  • AEB braking method, device and equipment facing pedestrian shielding scene and medium

    CN120621347A

  • Vehicle obstacle avoidance and path planning method based on intelligent driving

    CN121043907A