Planned trajectory generation method and apparatus, and electronic device and storage medium

By optimizing the POMDP model and Frenet coordinate system, the problem of not considering the uncertain behavior of traffic participants in the existing technology is solved, and a safe planning trajectory that conforms to driving habits is generated, which improves the decision-making efficiency and stability of autonomous driving.

WO2026157675A1PCT designated stage Publication Date: 2026-07-30SAIC GM WULING AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAIC GM WULING AUTOMOBILE CO LTD
Filing Date
2025-12-17
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing technologies do not adequately consider the uncertain behaviors of other traffic participants, resulting in an imperfect final planned trajectory. Furthermore, existing decision-making methods lack mechanisms to address uncertain behaviors, requiring adjustments to the planned trajectory during driving.

Method used

The behavior interaction between the autonomous vehicle and the adjacent vehicle is modeled using a partially observable Markov decision process (POMDP). The predicted trajectories of the adjacent vehicles are classified into risk trajectory sets and safe trajectory sets based on potential collision risks. The state space is reduced in dimensionality, and the planned trajectory is optimized by combining the Frenet coordinate system.

Benefits of technology

It improves the efficiency and stability of decision-making and planning, ensures vehicle driving safety, and generates planned trajectories that conform to actual driving habits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025143065_30072026_PF_FP_ABST
    Figure CN2025143065_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a planned trajectory generation method and apparatus, and an electronic device and a storage medium. The method comprises: acquiring a state space and an action space of an ego vehicle; on the basis of a potential collision risk and the action space, classifying predicted trajectories of a neighboring vehicle, so as to obtain a risky trajectory set and a safe trajectory set; and modeling the state space, the action space, the risky trajectory set and the safe trajectory set by means of a partially observable Markov decision process, so as to obtain a reference trajectory. The decision planning process under the uncertainty of neighboring vehicle behaviors is modeled as a POMDP, such that uncertain behaviors of the neighboring vehicle are fully predicted, thereby ensuring the driving safety of an ego vehicle. In addition, considering that the dimensionality of a predicted trajectory cluster for multiple traffic participants in complex scenarios is excessively high, the POMDP model is prone to the curse of dimensionality. Therefore, predicted trajectories of a neighboring vehicle are classified into a risky trajectory set and a safe trajectory set by means of a potential collision risk, thereby achieving state space dimensionality reduction, and effectively improving decision planning efficiency and decision stability.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, devices, electronic equipment and storage media for planning trajectory generation

[0001] This application claims priority to Chinese patent application filed on January 23, 2025, with application number 202510113155.4 and entitled "Planning Trajectory Generation Method, Apparatus, Electronic Device and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of intelligent driving technology, specifically to a method, apparatus, electronic device, and storage medium for generating a planned trajectory. Background Technology

[0003] With the rapid development of automotive technology, more and more cars are equipped with autonomous driving technology. The decision planning module is one of the key parts of autonomous driving. The decision planning module is a process of making some purposeful decisions for a certain goal. It usually refers to getting from the starting point to the destination while avoiding obstacles, and continuously optimizing the planned trajectory and behavior to ensure the safety and comfort of passengers.

[0004] Some related technologies directly sample trajectory clusters from the planning space and then evaluate the final planned trajectory using indicators such as safety and comfort. However, they do not consider the behavioral interactions with other traffic participants, which may result in an imperfect final planned trajectory, requiring adjustments during driving.

[0005] In addition, some related technologies, such as decision-making methods based on optimization or game theory, although modeling the dynamics of behavioral interactions, only consider the deterministic prediction of the future movement state of surrounding traffic participants and lack a mechanism to deal with uncertain behaviors.

[0006] It should be noted that the information disclosed in the background section of this application is intended only to enhance the understanding of the general background of this application, and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention

[0007] In view of this, this application provides a method, apparatus, electronic device and storage medium for generating a planned trajectory, in order to solve the problem that the final planned trajectory is not perfect due to the lack of sufficient consideration of the uncertain behavior of other traffic participants in the prior art.

[0008] In a first aspect, embodiments of this application provide a method for generating a planned trajectory, including:

[0009] Acquire the state space and the action space of the vehicle. The state space includes the position information and motion state information of the vehicle and the adjacent vehicles. The action space includes the driving actions of the vehicle.

[0010] Based on the potential collision risk and the action space, the predicted trajectories of the adjacent vehicles are classified to obtain a set of risk trajectories and a set of safe trajectories.

[0011] By using a partially observable Markov decision process, the state space, the action space, the set of risk trajectories, and the set of safe trajectories are modeled to obtain a reference trajectory.

[0012] In this embodiment, the decision-making process under uncertainty regarding the behavior of other vehicles is modeled as a Partially Observable Markov Decision Process (POMDP), thereby fully predicting the uncertain behavior of other vehicles and ensuring the driving safety of the vehicle itself. Furthermore, considering the high dimensionality of predicted trajectory clusters from multiple traffic participants in complex scenarios, which can easily lead to the curse of dimensionality in POMDP decision-making models, the predicted trajectories of other vehicles are classified into risk trajectory sets and safe trajectory sets based on potential collision risks. This achieves dimensionality reduction of the state space, effectively improving decision-making efficiency and stability.

[0013] In one possible implementation, the vehicle's motion space includes a lateral behavior library and a longitudinal behavior library, and obtaining the vehicle's motion space includes:

[0014] Based on the global planning and lane-level map, obtain the set of feasible lanes;

[0015] Extract the centerline of the lane from the set of feasible lanes to obtain a set of reference paths, which is the lateral behavior library.

[0016] Obtain the acceleration set, which is the longitudinal behavior library.

[0017] In this embodiment, a set of feasible lanes is obtained based on global planning and a lane-level map. Lane centerlines are then extracted from these lanes to obtain a reference path set, which is the lateral behavior library. The acceleration set is used as the longitudinal behavior library according to typical longitudinal motion patterns. It is understood that the lateral behavior library in this application uses the reference path set, not the lane set. Lane centerlines, as paths, are easier to calculate than lanes, which helps reduce the algorithm's burden.

[0018] In one possible implementation, before classifying the predicted trajectories of the adjacent vehicles based on the potential collision risk and the action space, the method further includes:

[0019] Probability calculations are performed on the predicted trajectories of the adjacent vehicles to obtain the probability distribution of the predicted trajectories;

[0020] Based on the predicted trajectory probability distribution, predicted trajectories with a probability less than or equal to a preset probability threshold are deleted.

[0021] In this embodiment, before processing the predicted trajectory of the adjacent vehicle, a probability calculation is performed on the predicted trajectory to obtain its probability distribution. Based on this probability distribution, predicted trajectories with probabilities less than or equal to a preset probability threshold are deleted. It is understood that pre-calculating the probability of the predicted trajectory allows for the early removal of trajectories that are almost impossible to occur, thus avoiding any impact on decision-making safety and significantly optimizing real-time performance.

[0022] In one possible implementation, classifying the predicted trajectories of adjacent vehicles based on potential collision risks and the action space to obtain a set of risk trajectories and a set of safe trajectories includes:

[0023] The vehicle's driving lane is determined based on the aforementioned action space;

[0024] Determine whether the predicted trajectory of the adjacent vehicle intersects with the driving lane of the vehicle;

[0025] If the predicted trajectory of any of the adjacent vehicles intersects with the driving lane of the vehicle, then the predicted trajectory of any of the adjacent vehicles is determined to be a risk trajectory.

[0026] If the predicted trajectory of any of the adjacent vehicles does not intersect with the driving lane of the vehicle, then the predicted trajectory of any of the adjacent vehicles is determined to be a safe trajectory.

[0027] In this embodiment, the predicted trajectory of the adjacent vehicle is determined to be either a risky or safe trajectory by judging whether there is an intersection between the predicted trajectory of the adjacent vehicle and the driving lane of the own vehicle. This can be understood as transforming the probability distribution of the predicted trajectory into a probability distribution of whether there is a potential collision risk with the own vehicle. The confidence state space is directly reduced to two states, reducing the solution time of POMDP and improving real-time efficiency.

[0028] In one possible implementation, the reference trajectory is a trajectory generated with reference to the lane centerline, and after obtaining the reference trajectory, the method further includes:

[0029] The reference trajectory is adjusted laterally to obtain the planned trajectory.

[0030] In this embodiment of the application, the reference trajectory is a trajectory generated with the lane center line as a reference. However, when a vehicle is driving, it obviously cannot drive along the lane center line, so it is necessary to adjust the reference trajectory laterally to obtain a planned trajectory that is more in line with actual driving habits.

[0031] In one possible implementation, the lateral adjustment of the reference trajectory to obtain the planned trajectory includes:

[0032] Construct a Frenet coordinate system using the tangent and normal vectors of the reference trajectory;

[0033] In the Frenet coordinate system, a planned trajectory is generated at a position at a preset distance from the reference trajectory.

[0034] In this embodiment, a planned trajectory is generated at a preset distance from the reference trajectory using the Frenet coordinate system. It is understood that the Frenet coordinate system can more intuitively represent the vehicle's position on a curved road, making it more conducive to generating a planned trajectory that meets expectations.

[0035] In one possible implementation, the step of modeling the state space, the action space, the set of risk trajectories, and the set of safe trajectories through a partially observable Markov decision process to obtain a reference trajectory includes:

[0036] By using a partially observable Markov decision process, the state space, the action space, the set of risk trajectories, and the set of safe trajectories are modeled to obtain a set of feasible trajectories.

[0037] A reference trajectory is obtained from the set of feasible trajectories using a reward function.

[0038] In this embodiment, after POMDP modeling, a set of feasible trajectories is obtained. A reference trajectory is then obtained from this set using a reward function. It is understood that the set of feasible trajectories includes multiple trajectories, and the trajectory generation device needs to select a safe, reasonable trajectory that ensures passenger comfort and high driving efficiency as the reference trajectory.

[0039] Secondly, embodiments of this application provide a planning trajectory generation device, comprising:

[0040] The information acquisition module is used to acquire the state space and the action space of the vehicle. The state space includes the position information and motion state information of the vehicle and the adjacent vehicles. The action space includes the driving action of the vehicle.

[0041] The predicted trajectory classification module is used to classify the predicted trajectories of the adjacent vehicles based on the potential collision risk and the action space, and obtain a risk trajectory set and a safe trajectory set.

[0042] The modeling module is used to model the state space, the action space, the risk trajectory set, and the safe trajectory set through a partially observable Markov decision process to obtain a reference trajectory.

[0043] Thirdly, embodiments of this application provide an electronic device, including:

[0044] processor;

[0045] Memory;

[0046] And a computer program, wherein the computer program is stored in the memory, the computer program including instructions that, when executed by the processor, cause the electronic device to perform the method described in any one of the first aspects.

[0047] Fourthly, embodiments of this application provide a computer-readable storage medium, characterized in that the computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform the method described in any one of the first aspects.

[0048] It is understood that the trajectory generation apparatus provided in the second aspect, the electronic device provided in the third aspect, and the computer-readable storage medium provided in the fourth aspect are used to execute the method provided in this application. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description

[0049] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 is a flowchart illustrating a method for generating a planned trajectory according to an embodiment of this application;

[0051] Figure 2 is a lane diagram of a lateral motion library provided in an embodiment of this application;

[0052] Figure 3 is a schematic diagram of a potential collision risk assessment provided by an embodiment of this application;

[0053] Figure 4 is a schematic diagram of probability calculation for a representative trajectory provided in an embodiment of this application;

[0054] Figure 5 is a schematic diagram of a vehicle lane-changing trajectory provided in an embodiment of this application;

[0055] Figure 6 is a schematic diagram of the architecture of another trajectory generation method provided in the embodiments of this application;

[0056] Figure 7 is a schematic diagram of a planning trajectory generation device provided in an embodiment of this application;

[0057] Figure 8 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0058] To better understand the technical solution of this application, the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0059] It should be understood that the described embodiments are merely some, not all, of the embodiments in this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0060] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0061] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0062] With the rapid development of automotive technology, more and more cars are equipped with autonomous driving technology. The decision planning module is one of the key parts of autonomous driving. The decision planning module is a process of making some purposeful decisions for a certain goal. It usually refers to getting from the starting point to the destination while avoiding obstacles, and continuously optimizing the planned trajectory and behavior to ensure the safety and comfort of passengers.

[0063] Some related technologies directly sample trajectory clusters from the planning space and then evaluate the final planned trajectory using indicators such as safety and comfort. However, they do not consider the behavioral interactions with other traffic participants, which may result in an imperfect final planned trajectory, requiring adjustments during driving.

[0064] In addition, some related technologies, such as decision-making methods based on optimization or game theory, although modeling the dynamics of behavioral interactions, only consider the deterministic prediction of the future movement state of surrounding traffic participants and lack a mechanism to deal with uncertain behaviors.

[0065] To address the aforementioned issues, this application provides a trajectory generation method that models the decision-making process under uncertainty regarding the behavior of other vehicles as a Partially Observable Markov Decision Process (POMDP), thereby fully predicting the uncertain behavior of other vehicles and ensuring the driving safety of the vehicle itself. Furthermore, considering the excessively high dimensionality of predicted trajectory clusters from multiple traffic participants in complex scenarios, which can easily lead to the curse of dimensionality in POMDP decision-making models, this application categorizes the predicted trajectories of other vehicles into risk trajectory sets and safe trajectory sets based on potential collision risks. This achieves dimensionality reduction of the state space, effectively improving decision-making efficiency and stability. The following detailed description, in conjunction with the accompanying drawings and specific embodiments, further illustrates this method.

[0066] Referring to Figure 1, it is a flowchart illustrating a trajectory generation method provided in an embodiment of this application. As shown in Figure 1, it mainly includes the following steps.

[0067] Step S101: Obtain the state space and the motion space of the vehicle.

[0068] The state space includes the position and motion information of the vehicle and the adjacent vehicle. Specifically, it includes the vehicle's and adjacent vehicle's positions (x, y), driving direction θ, speed v, distance traveled along the target reference path s, and the adjacent vehicle's driving intention I. Driving intention I characterizes the adjacent vehicle's future trajectory and is a part of the system state that cannot be directly observed. In this embodiment, it is assumed that the adjacent vehicle's driving intention I is a trajectory from the predicted trajectory cluster output by the prediction model, and is uncertain during decision-making. Monte Carlo sampling is used to examine the impact of this uncertainty on collision risk assessment. The POMDP decision-making and programming model needs to find the optimal strategy that satisfies the maximum discounted cumulative reward under the probability distribution of driving intentions. Other road information and fixed buildings are treated as static references and are not included in the state space.

[0069] The state definition in this application embodiment is as follows: s = [s ego s1, s2, ..., s n ]s ego =[x,y,θ,v]s i =[x,y,θ,v,I]

[0070] Among them, s ego Indicates the status of the vehicle, s i ,i∈{1,2,......,n} represents the state of the vehicles surrounding the vehicle, and n is the number of vehicles.

[0071] In the embodiments of this application, the action space includes vehicle driving actions, specifically, the action space includes two parts: a longitudinal action library and a lateral action library.

[0072] In this embodiment, the longitudinal behavior set includes four discrete accelerations:

[0073] Among them, A lon For the longitudinal movement of the vehicle, it can be understood that when A... lon for At that time, the acceleration is 0 m / s². 2 At this time, the vehicle is moving at a constant speed; when A lon for At that time, the acceleration was 1.2 m / s². 2 At this moment, the vehicle accelerates; when A lon for At that time, the acceleration was -1.8 m / s². 2 At this time, the vehicle slows down; when A lon for At that time, the acceleration was -3.6 m / s². 2 At this point, the vehicle is forced to move.

[0074] In some possible implementations, if the situation ahead is too urgent and a collision cannot be avoided even with the acceleration corresponding to the forced longitudinal action described above, then an emergency braking mode is entered and POMDP replanning is performed using the following longitudinal action:

[0075] By using the emergency braking mode, the safety of emergency braking is ensured, while avoiding the continuous expansion of the action space, which would affect the POMDP solution efficiency.

[0076] In some related technologies, only longitudinal velocity planning is performed, while lateral trajectory planning is carried out using other methods outside the POMDP model. In other related technologies, although both longitudinal velocity and lateral trajectory planning are modeled in the POMDP model, the lateral trajectory is fixedly defined as driving according to a predetermined pattern, including left-turn lanes, lane keeping, right lane changes, straight ahead at intersections, left turns at intersections, and right turns at intersections. This method of predefining driving actions based on traffic scenarios may not be able to traverse all necessary traffic scenarios, and these redundant actions increase or decrease the dimensionality of the POMDP action space, directly leading to an exponential increase in the time complexity of POMDP solution.

[0077] To address this issue, in this embodiment, the lateral behavior library is dynamically adjusted to the set of currently drivable reference paths, and its action space is more consistent with real-time traffic scenarios.

[0078] Specifically, guided by a globally planned path in the lane-level map, the currently drivable lane is determined and its centerline is read. Then, a reference path is generated using piecewise cubic B-spline curves.

[0079] First, the aiming distance within the decision period is determined based on the current speed and maximum acceleration, using the following formula:

[0080] In the formula, It is the current speed of the vehicle, a max,acc This represents the maximum permissible acceleration, and T is the duration of the decision-making time domain. Then, path points within the pre-aimed distance between the current lane and the target lane are selected as spline curve control points P. i (i = 0, 1, ..., n), each small segment of the cubic B-spline curve consists of 4 control points P. i P i+1 P i+2 and P i+3 The generated piecewise cubic B-spline curve can be described as follows:

[0081] Finally, the generated lateral motion library A lat ={path 0 ,path 1 ,……,path n}, where path i (i = 0, 1, ..., n) represents the available lanes for the vehicle. For ease of understanding, this application uses two available lanes as an example to provide a detailed description of the available lanes.

[0082] Referring to Figure 2, this is a lane diagram of a lateral action library provided in an embodiment of this application. As shown in Figure 2, a vehicle is traveling straight in one lane of a two-lane road system. In this embodiment, the first path in the lateral action library is... 0 To travel straight along the current lane, use the second path in the lateral action library. 1 To change lanes to the left.

[0083] The action space A of the POMDP model can be obtained by combining elements from the vertical action library and the horizontal action library in pairs, as shown in the following formula:

[0084] It should be noted that the vertical actions in the formulas of this application's embodiments include and However, in practical applications, longitudinal movements may include other movements, and this application embodiment does not impose specific limitations on them.

[0085] Furthermore, in most cases, there are no more than three drivable paths: the current lane and the adjacent lanes on the left and right. Therefore, the dimension of the lateral action library mostly changes dynamically within the range of 0-3, resulting in the dimension of the corresponding action space A mostly changing dynamically within the range of 0-12. Moreover, to avoid vehicles changing lanes back and forth within a decision cycle, this embodiment limits the lateral path change operation to only once per decision cycle, thereby saving solution time and improving work efficiency.

[0086] Step S102: Based on the potential collision risk and the action space, classify the predicted trajectories of the adjacent vehicles to obtain a set of risk trajectories and a set of safe trajectories.

[0087] Specifically, based on the action space, the driving lane of the vehicle is determined, and it is determined whether the predicted trajectory of the adjacent vehicle intersects with the driving lane of the vehicle. If the predicted trajectory of any adjacent vehicle intersects with the driving lane of the vehicle, the predicted trajectory is determined to be a risky trajectory; if the predicted trajectory of any adjacent vehicle does not intersect with the driving trajectory of the vehicle, the predicted trajectory is determined to be a safe trajectory.

[0088] For ease of understanding, this application provides a schematic diagram for assessing potential collision risks.

[0089] [Corrected according to Rule 91, 05.02.2026] Refer to Figure 3, which is a schematic diagram of a potential collision risk assessment provided by an embodiment of this application. As shown in Figure 3, adjacent vehicles A and B are adjacent vehicles, and vehicle M is the vehicle itself. Taking a two-lane road as an example, vehicle M has two driving decision candidates: (a) lane keeping and (b) lane changing. In each decision, the centerline of the target lane is used as the local planning reference path, and then the existence of a potential collision risk is determined by judging whether the predicted trajectory cluster intersects with the lane area occupied by the reference path. As shown in Figure 3, the reference path is a bold dashed line, and trajectory C in the predicted trajectory cluster of the adjacent vehicle is... safe For a safe trajectory, trajectory C risk For risk trajectory.

[0090] [Corrected according to Rule 91, 05.02.2026] As shown in Figure 3(a), vehicle M is currently traveling in the right lane. When the vehicle maintains its current lane, the trajectories in the predicted trajectory clusters of vehicles A and B that intersect with the right lane are all risky trajectories, while the trajectories in the predicted trajectory clusters of vehicles A and B that do not intersect with the right lane are all safe trajectories. Similarly, as shown in Figure 3(b), vehicle M is currently traveling in the right lane, but the vehicle's driving decision is to change lanes. Therefore, the vehicle's corresponding reference path is the center line of the left lane. Since all trajectories in the predicted trajectory clusters of vehicles A and B intersect with the left lane, all predicted trajectories of vehicles A and B are risky trajectories.

[0091] It is understandable that the above classification of potential collision risks based on the predicted trajectories of adjacent vehicles is static and unrelated to the future speed of the vehicle itself. This is because the vehicle's speed is obtained by the POMDP decision model and is not known at this point. However, in this embodiment, if a predicted trajectory is determined to be a safe trajectory, it means that the vehicle will not collide with adjacent vehicles regardless of its speed.

[0092] After classifying the potential collision risk of the predicted trajectories of the adjacent vehicles, the predicted trajectories of the adjacent vehicles are divided into two main categories according to whether a certain candidate strategy of the vehicle itself poses a potential collision hazard: Risk Trajectory C risk and safe trajectory C safe And calculate the sum of probabilities for the corresponding categories. Then, respectively in C risk and C safe The trajectory with the highest probability is selected as the representative trajectory for collision safety. Specifically, the calculation method is as follows:

[0093] Where P represents probability, and track is the representative trajectory of the category, used for sampling and as the future trajectory for constructing the state transition model of surrounding vehicle i. For ease of understanding, this embodiment of the application uses the example of the vehicle traveling straight in the right lane and the presence of vehicle B in the left lane in Figure 3 to explain in detail the probability calculation of the representative trajectory.

[0094] [Correction 05.02.2026 based on Rule 91] Referring to Figure 4, this is a schematic diagram of probability calculation for a representative trajectory provided in an embodiment of this application. As shown in Figure 4, the predicted trajectory cluster of the adjacent vehicle B includes 495 trajectories, namely the safe trajectory C. safe and risk trajectory C risk The total number is 495, of which the highest probability safe trajectory (i.e., safe trajectory C) safe The representative trajectory track safe ) indicates a bold line. s The highest probability risk trajectory (i.e., risk trajectory C) risk The representative trajectory track risk ) indicates a bold line. r The track can be obtained through calculation. safe The predicted probability is 0.137, track risk The predicted probability is 0.863. The degree of threat posed by other vehicles to the vehicle can be inferred from the predicted probabilities corresponding to safe and risky trajectories.

[0095] It is understandable that there are usually dozens or even hundreds of predicted trajectory clusters for the vehicles beside the other vehicle. If these are directly input into the POMDP model as the future state of the vehicles beside the other vehicle to participate in collision detection, the computational load of the POMDP model will be very large. The predicted trajectory classification method provided in this application greatly reduces the solution complexity and transforms the confidence space of dozens or hundreds of predicted trajectory clusters into a confidence space with only two states: whether there is a collision risk or not, thereby improving work efficiency.

[0096] As mentioned above, there are hundreds or even thousands of predicted trajectories for the vehicles passing by, many of which are extremely low-probability trajectories. Treating these extremely low-probability trajectories equally and including them in subsequent calculations is meaningless and would severely degrade the real-time performance of the decision-making system.

[0097] Therefore, in one possible implementation, before classifying the predicted trajectories of the adjacent vehicles, the method further includes calculating the probability of the predicted trajectories to obtain a predicted trajectory probability distribution; based on the predicted trajectory probability distribution, predicted trajectories with probabilities less than or equal to a preset probability threshold are deleted. The preset probability threshold is a fixed threshold set based on debugging experience; those skilled in the art can set it to any value according to actual needs, and this application embodiment does not impose specific limitations on it.

[0098] Understandably, extremely low-probability trajectories—those that are almost impossible to occur—are as follows: for example, the historical trajectory of a neighboring vehicle consistently remains straight within the current lane. Therefore, in the predicted trajectory clusters output by the prediction module, the probability of left turns is extremely low, indicating that left turns are very rare events. Removing these trajectories would not only not affect decision-making safety but would also significantly optimize real-time performance.

[0099] Step S103: Model the state space, action space, risk trajectory set, and safe trajectory set through a partially observable Markov decision process to obtain a reference trajectory.

[0100] Specifically, in order to obtain the states of the driver and other vehicles, this embodiment uses a state transition model to describe the evolution of the states of each participant in a traffic scenario over time. The state transition model is a mathematical model of the state transition probability p(s′|s, a). This embodiment involves the behavioral interactions between the driver and other vehicles, so the state transition model is explained from the perspectives of both the driver and other vehicles.

[0101] This application primarily studies the decision-making problem under the uncertainty of the driving intention of a neighboring vehicle. It assumes that the position, speed, and heading of both the vehicle and the neighboring vehicle are accurate, with errors small enough not to affect the decision result. Only the driving intention of the neighboring vehicle is uncertain. Furthermore, since the decision-making module is more abstract than the control module, its accuracy requirements for positioning and state evolution are not very high. Therefore, the state transition model of the vehicle can be completely determined by the vehicle's kinematics model, without involving a complex dynamic model. From the definition of the action space, it can be seen that the vehicle's decision output first selects a dynamically generated drivable path outside the model, and then performs longitudinal movement on the selected lateral path. Therefore, the vehicle's state transition model can be expressed as:

[0102] In the formula, a lon It is the acceleration corresponding to the longitudinal action taken at a certain decision step, path i S represents the drivable path i corresponding to the lateral action taken in this decision step, where S is the path along the path. i The distance traveled, initially set to path. i The starting point and the current point of the vehicle are on the path i The distance between the projections, x, y, θ, lies on the selected path. i The above is obtained by interpolation using the current travel distance S, where △t is the decision step size.

[0103] The state transition model of the vehicle beside the driver differs from that of the vehicle itself. The state update of the vehicle is performed along the selected lateral path. i The vehicle i is moving longitudinally, while surrounding vehicles do not involve changes in driving actions, and their predicted trajectories are a series of [x,y] trajectory points that change over time. Therefore, their state updates are obtained directly by interpolating the predicted trajectory at each decision step. However, the driving intention I represented by the predicted trajectory cannot be obtained as a true value; it can only be evaluated and decided upon based on its probability distribution. Therefore, this embodiment randomly samples a cluster of predicted trajectories from the probability distribution of driving intention I and uses it as the future trajectory of the vehicle i in a certain decision-making process. Calculating the average of the discounted cumulative reward after multiple Monte Carlo sampling simulations can approximately approximate the optimal decision under the probability distribution of driving intention I. Therefore, in a certain Monte Carlo sampling simulation, driving intention I is unique and deterministic, i.e., the predicted trajectory sampled in this simulation. Thus, the state transition model of the vehicle i can be represented as the state transition of the vehicle i on the sampled predicted trajectory track. pred Above, directly from the current decision step time t ′ Interpolation yields new trajectory points:

[0104] In the formula, track safe For the safe trajectory mentioned above, track riskThis refers to the risk trajectory described above.

[0105] Furthermore, in this embodiment, the posterior probability distribution of the current state is estimated using the observation z, thereby obtaining an observation space Z similar to the state space, which is defined as: Z = [Z ego Z1, Z2, ..., Z n Z ego = [x,y,θ,v]Z i = [x, y, θ, v, I]

[0106] Among them, z ego This represents the vehicle's measurements, including its position, heading angle, and speed information, z. i The measurement values ​​of the adjacent vehicle are represented, including the adjacent vehicle's position, heading angle, speed, and driving intention. The driving intention I of the adjacent vehicle is the probability distribution of the future trajectory after the predicted trajectory cluster is reduced in dimensionality by clustering with potential collision risks.

[0107] The observation model O represents the probability p(z|s′,a) of the measured value when the vehicle reaches the new state s′ after taking action a in the current state. The sensing system in this embodiment is based on high-precision combined inertial navigation positioning and high-precision lidar. Both the vehicle and the adjacent vehicle have centimeter-level positioning accuracy. The sensing error level is far less than the impact of the uncertainty of the adjacent vehicle's intention on the decision-making result. Therefore, this embodiment does not consider observation errors and assumes that the measured value is equivalent to its corresponding state value. That is, after executing action a, the vehicle's measured value changes completely with the state transition model. The adjacent vehicle's state and measured value are determined by its future predicted trajectory. Although its driving intention cannot be directly observed, its predicted trajectory probability distribution can be output by the prediction module. Then, within the decision-making period, after fully considering its probability distribution and uncertainty, the optimal future trajectory is sampled as the basis for updating the adjacent vehicle's state and measured value.

[0108] In one possible implementation, in order to obtain a reference trajectory with greater benefits, the state space, action space, risk trajectory set, and safe trajectory set are modeled using POMDP. After obtaining the feasible trajectory set, the reference trajectory is obtained from the feasible trajectory set through a reward function.

[0109] Specifically, the reward function is the immediate reward after performing action a in confidence state b, and it is generally quantified and compared using multiple evaluation metrics to assess the merits of action a in state b. The optimal strategy π is to select the sequence of state actions with the highest cumulative discounted reward by accumulating the immediate rewards over the entire decision-making cycle. This application's embodiments consider the global task completion degree R... goal Safety R safe Comfort R comfort And driving efficiency R vThe reward function was designed from several aspects, and the overall reward function and the effects of these aspects are: r(b, a) = R goal +R safe +R comfort +R v

[0110] The global task constraint is as follows: given the global path, the decision-seeking distance L is determined according to the formula. preview A local reference path can be extracted from the global reference path, and the endpoint of the local reference path can be used as the target point for the current decision cycle. Under the premise of ensuring driving safety and comfort, the further the distance traveled along the local reference path, the closer to the target point, the higher the task completion rate under the guidance of the global path, and the greater the benefit. The formula is as follows:

[0111] in, and S is the weighting coefficient. goal Let l be the distance traveled by the vehicle from the projection point of the local reference path to the target point. goal This is the lateral and vertical distance from the vehicle's actual position to the local reference path. Unreasonable lane-changing actions can be suppressed through global task constraints.

[0112] The specific driving safety constraint is as follows: Starting from the current moment, assess the potential safety of the vehicle colliding with another vehicle within the future decision-making cycle. The driving safety assessment is based on whether a collision occurs within a certain decision-step time interval Δt between the vehicle's planned trajectory and the other vehicle's predicted trajectory, using the following formula:

[0113] Among them, w safe w is the weighting coefficient, relative to other weights. safe A relatively large constant, such as 1000, is typically used to penalize actions that cause collisions. ego It is the vehicle speed at the moment of the collision, and the constant const is used to ensure the penalty when the vehicle speed is 0.

[0114] The comfort constraint specifically refers to the fact that occupant comfort is an important indicator for evaluating intelligent driving systems, and current research mainly uses acceleration and its rate of change to characterize it. Acceleration includes two dimensions: longitudinal and lateral acceleration.

[0115] Where v is the longitudinal vehicle speed, θ is the vehicle heading angle, and a lon For longitudinal acceleration, a lat For lateral acceleration, an excessively large rate of change of acceleration j manifests as frequent acceleration and deceleration at the operational level, leading to significant discomfort. When both acceleration and its rate of change exceed critical values, comfort deteriorates as both increase, as follows: R comf ort=-abs(w lon_a a lon +w lat_a a lat +w lon_j j lon +w lat_j j lat )

[0116] Among them, w lon_a w lat_a w lon_j and w lat_j These are the corresponding weighting coefficients.

[0117] The driving efficiency constraint is specifically set as follows: In order to encourage intelligent driving vehicles to travel at the desired speed while ensuring safety and comfort, a speed constraint is set as follows:

[0118] Among them, v max For the maximum speed permitted by traffic rules, w v These are the weighting coefficients.

[0119] It should be noted that all the weighting coefficients mentioned above are values ​​set based on experience. Those skilled in the art can set the weighting coefficients to any value according to actual needs, and this application embodiment does not impose any specific restrictions on this.

[0120] After constructing the POMDP decision-making and programming model and reducing its dimensionality in the state space to improve decision-making efficiency and stability, the POMDP model is solved using the Despot solver. The final output is a reference trajectory along the road centerline containing speed information. The speed is obtained by integrating the longitudinal action sequence, i.e., the desired acceleration. The reference path corresponding to the reference trajectory is a drivable path selected by the lateral action sequence, with its endpoint located only at the center of the corresponding lane.

[0121] Referring to Figure 5, a schematic diagram of a lane-changing trajectory provided by an embodiment of this application is shown. As shown in Figure 5, the endpoint of the reference trajectory is located at the center of the corresponding lane. The purpose of this is to reduce the POMDP action space dimension, but it brings a problem: the vehicle always travels along the lane centerline, which obviously does not conform to actual driving habits. Therefore, based on the reference trajectory, this embodiment of the application keeps the speed information unchanged, samples more planned trajectories laterally, and selects the optimal planned trajectory according to performance indicators, thereby improving the rationality of the planned trajectory.

[0122] Specifically, the reference trajectory output by POMDP is used as a reference line, and a Frenet coordinate system is constructed using the tangent vector and normal vector of this line. In the Frenet coordinate system, the horizontal coordinate 's' describes the distance along the reference line, and the vertical coordinate 'd' represents the displacement of the normal vector at the current position along the reference line, i.e., the distance deviating from the reference line. The Frenet coordinate system is more effective than the Cartesian coordinate system in describing how far a vehicle has traveled along a curved lane and whether it has deviated from the lane center. Keeping the reference trajectory distance constant, a set of end-point normal displacements 'd' of different magnitudes are sampled within the target lane. Combined with the starting state, a set of fifth-order polynomial curves spreading across the lateral space of the target lane can be obtained. It is worth noting that, in order not to compromise the existing collision safety of the reference trajectory, the sampled 'd' must satisfy the following formula; otherwise, the newly sampled planned trajectory set will need to undergo collision risk assessment under the uncertainty of the surrounding vehicles' intentions, thus rendering the reference trajectory obtained using POMDP meaningless.

[0123] Among them, width ego For the width of the vehicle, l margin For safety reasons, the vehicle body has a certain amount of lateral space. threshold This is the lateral safety distance threshold.

[0124] Unlike most current methods that use Frenet coordinate sampling for trajectory planning, the reference line here is not a global path, but a reference trajectory that has been modeled and solved using POMDP and has already satisfied indicators such as collision safety, driving efficiency, and comfort as much as possible. Here, we only perform fine-tuning of the reference trajectory. Therefore, we no longer sample the distance s along the reference line direction, but only sample the end-point normal displacement d. As a result, the generated candidate set of planned trajectories is very small, and the computational cost is not high.

[0125] After obtaining the above candidate trajectory set, the performance of each candidate trajectory needs to be calculated according to the cost evaluation function. Then, the optimal trajectory is selected as the planned trajectory output to the vehicle control module. Specifically:

[0126] In the formula, J{d,k,c} is the loss function term, w{d,k,c} is the corresponding weight coefficient, and i * The most valid trajectory index is N, which is the total number of candidate trajectories. Since the collision safety of the reference trajectory is guaranteed by POMDP modeling, and the sampling range of the lateral distance d is limited by the formula, there is no need to perform a collision risk assessment here.

[0127] Loss Item J d This indicates the degree of deviation from the reference trajectory, with the aim of guiding the vehicle to tend to travel along the center of the road in most situations. J d The specific calculation method is as follows:

[0128] In the formula, d i It is the sampling deviation distance of candidate trajectory i, d max It is the maximum deviation distance.

[0129] Trajectory curvature smoothing cost J k This is to improve ride comfort and avoid violent lateral movement shocks. J k The specific calculation method is as follows:

[0130] In the formula, τ i The candidate trajectory i is given, and k is the curvature of the trajectory points.

[0131] Loss Item J c This is to penalize the discontinuity between two planned trajectories, preventing vehicle instability such as swaying and overshooting. c The specific calculation method is as follows:

[0132] In the formula, d pre,i* It is the deviation distance at the end of the previously planned optimal trajectory.

[0133] Corresponding to the above embodiments, this application also provides another method for generating planning trajectories.

[0134] Referring to Figure 6, this is a schematic diagram of the architecture of another trajectory generation method provided in this application embodiment. As shown in Figure 6, the state space is obtained through the positioning system, the set of feasible lanes is obtained through global planning and lane-level map, the set of reference paths is obtained based on the set of feasible lanes, the set of reference paths can form a lateral behavior library, the set of accelerations is the longitudinal behavior library, and the lateral behavior library and the longitudinal behavior library form the action space.

[0135] The trajectory generation device acquires the predicted trajectory of the adjacent vehicle and calculates its probability distribution. The probability distribution is then pruned (predicted trajectories with probabilities less than or equal to a preset probability threshold are deleted). The predicted trajectories are categorized into risky and safe trajectories based on a set of reference paths. Specifically, the probability of each risky and safe trajectory is summed, and the maximum value among the risky and safe trajectories is used as the representative trajectory. Two representative trajectories are sampled probabilistically to obtain the adjacent vehicle's state transition. The safety distance is dynamically adjusted, and the collision risk is assessed based on the ST distribution diagram. The action corresponding to the maximum benefit is obtained through a reward function.

[0136] The state space, action space, vehicle state transition and reward function are modeled by POMDP, and the lateral and longitudinal behavior sequences are solved by DESPOT to generate a reference trajectory. The planned trajectory is obtained by lateral sampling and optimization of the reference trajectory.

[0137] Corresponding to the above embodiments, this application also provides a planning trajectory generation device.

[0138] Referring to Figure 7, a schematic diagram of a trajectory planning generation device provided in an embodiment of this application is shown. As shown in Figure 7, the trajectory planning generation device includes: an information acquisition module 701, a trajectory prediction and classification module 702, and a modeling module 703. These components communicate through one or more buses. Those skilled in the art will understand that the structure of the electronic device shown in the figure does not constitute a limitation on the embodiment of this application. It can be a bus-shaped structure or a star-shaped structure, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0139] Among them, the information acquisition module 701 is used to acquire the state space and the action space of the vehicle. The state space includes the position information and motion state information of the vehicle and the adjacent vehicles, and the action space includes the driving action of the vehicle.

[0140] The predicted trajectory classification module 702 is used to classify the predicted trajectories of the adjacent vehicles according to the potential collision risk and the action space, and obtain a risk trajectory set and a safe trajectory set.

[0141] The modeling module 703 is used to model the state space, the action space, the risk trajectory set, and the safe trajectory set through a partially observable Markov decision process to obtain a reference trajectory.

[0142] Corresponding to the above embodiments, this application also provides an electronic device.

[0143] Referring to Figure 8, a schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. As shown in Figure 8, the electronic device 800 may include a processor 801, a memory 802, and a communication unit 803. These components communicate through one or more buses. Those skilled in the art will understand that the structure of the electronic device shown in the figure does not constitute a limitation on the embodiment of this application. It may be a bus topology or a star topology, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0144] The communication unit 803 is used to establish a communication channel, enabling the electronic device to communicate with other devices. It can receive user data sent by other devices or send user data to other devices.

[0145] The processor 801 serves as the control center of the electronic device, connecting various parts of the device via interfaces and lines. It executes software programs, instructions, and / or modules stored in the memory 802, and calls data stored in the memory to perform various functions and / or process data. The processor may be composed of integrated circuits (ICs), such as a single packaged IC or multiple packaged ICs with the same or different functions connected together. For example, the processor 801 may consist only of a central processing unit (CPU). In this embodiment, the CPU may have a single processing core or include multiple processing cores.

[0146] The memory 802 is used to store the execution instructions of the processor 801. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0147] When the execution instructions in memory 802 are executed by processor 801, the electronic device 800 is able to perform some or all of the steps in the embodiment shown in FIG1.

[0148] In a specific implementation, this application embodiment also provides a computer storage medium, wherein the computer storage medium may store a program, and when the program is executed, it may include some or all of the steps of the simulation scene generation method provided in various embodiments of this application. The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0149] In a specific implementation, this application also provides a computer program product, wherein the computer program product includes executable instructions, which, when executed on a computer, cause the computer to perform some or all of the steps in the various embodiments of the simulation scene generation method provided in this application.

[0150] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, or the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0151] Those skilled in the art will recognize that the units and algorithm steps described in the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of electronic hardware and software. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0152] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0153] In the several embodiments provided in this application, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0154] The same or similar parts between the various embodiments in this specification can be referred to mutually. In particular, the device embodiments and terminal embodiments are basically similar to the method embodiments, so the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.

Claims

1. A method for generating a planned trajectory, characterized in that, include: Acquire the state space and the action space of the vehicle. The state space includes the position information and motion state information of the vehicle and the adjacent vehicles. The action space includes the driving actions of the vehicle. Based on the potential collision risk and the action space, the predicted trajectories of the adjacent vehicles are classified to obtain a set of risk trajectories and a set of safe trajectories. By using a partially observable Markov decision process, the state space, the action space, the set of risk trajectories, and the set of safe trajectories are modeled to obtain a reference trajectory.

2. The method according to claim 1, characterized in that, The vehicle's motion space includes a lateral behavior library and a longitudinal behavior library. Obtaining the vehicle's motion space includes: Based on the global planning and lane-level map, obtain the set of feasible lanes; Extract the centerline of the lane from the set of feasible lanes to obtain a set of reference paths, which is the lateral behavior library. Obtain the acceleration set, which is the longitudinal behavior library.

3. The method according to claim 1, characterized in that, Before classifying the predicted trajectories of adjacent vehicles based on potential collision risks and the action space, the method further includes: Probability calculations are performed on the predicted trajectories of the adjacent vehicles to obtain the probability distribution of the predicted trajectories; Based on the predicted trajectory probability distribution, predicted trajectories with a probability less than or equal to a preset probability threshold are deleted.

4. The method according to claim 1, characterized in that, The process of classifying the predicted trajectories of adjacent vehicles based on potential collision risks and the action space to obtain a risk trajectory set and a safe trajectory set includes: The vehicle's driving lane is determined based on the aforementioned action space; Determine whether the predicted trajectory of the adjacent vehicle intersects with the driving lane of the vehicle; If the predicted trajectory of any of the adjacent vehicles intersects with the driving lane of the vehicle, then the predicted trajectory of any of the adjacent vehicles is determined to be a risk trajectory. If the predicted trajectory of any of the adjacent vehicles does not intersect with the driving lane of the vehicle, then the predicted trajectory of any of the adjacent vehicles is determined to be a safe trajectory.

5. The method according to claim 1, characterized in that, The reference trajectory is a trajectory generated with the lane centerline as a reference. After obtaining the reference trajectory, the following steps are also included: The reference trajectory is adjusted laterally to obtain the planned trajectory.

6. The method according to claim 5, characterized in that, The lateral adjustment of the reference trajectory to obtain the planned trajectory includes: Construct a Frenet coordinate system using the tangent and normal vectors of the reference trajectory; In the Frenet coordinate system, a planned trajectory is generated at a position at a preset distance from the reference trajectory.

7. The method according to claim 1, characterized in that, The process of modeling the state space, action space, risk trajectory set, and safe trajectory set through a partially observable Markov decision process to obtain a reference trajectory includes: By using a partially observable Markov decision process, the state space, the action space, the set of risk trajectories, and the set of safe trajectories are modeled to obtain a set of feasible trajectories. A reference trajectory is obtained from the set of feasible trajectories using a reward function.

8. A planning trajectory generation device, characterized in that, include: The information acquisition module is used to acquire the state space and the action space of the vehicle. The state space includes the position information and motion state information of the vehicle and the adjacent vehicles. The action space includes the driving action of the vehicle. The predicted trajectory classification module is used to classify the predicted trajectories of the adjacent vehicles based on the potential collision risk and the action space, and obtain a risk trajectory set and a safe trajectory set. The modeling module is used to model the state space, the action space, the risk trajectory set, and the safe trajectory set through a partially observable Markov decision process to obtain a reference trajectory.

9. An electronic device, characterized in that, include: processor; Memory; And a computer program, wherein the computer program is stored in the memory, the computer program including instructions that, when executed by the processor, cause the electronic device to perform the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 7.