Vehicle trajectory planning method, device and electronic equipment

CN122481789BActive Publication Date: 2026-09-22SHANGHAI JUNZHENG NETWORK TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610969873.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-22
Estimated Expiration
2046-07-01

AI Technical Summary

Technical Problem

[0004]但是,在实际复杂驾驶场景中,车辆的运行环境是不断变化的

Benefits of technology

[0024]本说明书实施例的车辆行驶轨迹规划方法,可以根据车辆行驶数据,确定惩罚权重的配置模式;可以根据所述配置模式,确定惩罚权重的取值;可以根据惩罚权重的取值,构建目标函数,所述目标函数用于通过多种优化维度的多个分段惩罚函数度量行驶轨迹的优劣,所述分段惩罚函数的多个函数分支对应多种惩罚方式;可以利用求解器对所述目标函数进行优化求解,得到行驶轨迹。本说明书实施例可以根据车辆行驶数据,确定惩罚权重的配置模式,可以根据所述配置模式,确定惩罚权重的取值。由此实现惩罚权重的取值与车辆行驶状态的动态适配。另外,目标函数可以通过多种优化维度的多个分段惩罚函数评估行驶轨迹的优劣,所述分段惩罚函数的多个函数分支对应多种惩罚方式,便于根据车辆行驶数据选择适用的函数分支。由此实现了惩罚方式与车辆行驶状态的动态适配。本说明书实施例通过惩罚权重取值与惩罚方式的环境动态适配,有效避免了因惩罚权重取值固定和惩罚方式单一导致的轨迹规划不合理问题,显著提高了车辆自动驾驶系统在复杂道路环境下的适应能力和运行可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122481789B_ABST
    Figure CN122481789B_ABST
Patent Text Reader

Abstract

Embodiments of the present specification disclose a vehicle driving trajectory planning method and device and electronic equipment. The method comprises: determining a configuration mode of a penalty weight according to vehicle driving data; determining a value of the penalty weight according to the configuration mode; constructing a target function according to the value of the penalty weight, the target function being used to measure the pros and cons of a driving trajectory through a plurality of segmented penalty functions of multiple optimization dimensions, a plurality of function branches of the segmented penalty function corresponding to multiple penalty modes; and obtaining a driving trajectory by optimizing and solving the target function by using a solver. Embodiments of the present specification can adapt to dynamically changing vehicle driving states, and improve the safety and rationality of the planned vehicle driving trajectory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of autonomous driving technology, and in particular to a vehicle trajectory planning method, apparatus, and electronic device. Background Technology

[0002] As autonomous driving technology evolves to higher levels, the requirements for vehicle adaptability to complex road environments are becoming increasingly stringent. Vehicle trajectory planning algorithms, as the core decision-making unit of autonomous driving systems, need to quickly generate safe, efficient, and smooth vehicle trajectories in dynamic scenarios. Their optimization performance directly determines the reliability and safety of the autonomous driving system.

[0003] Vehicle trajectory planning algorithms optimize vehicle trajectories across multiple objectives through an objective function. This objective function includes penalty weights, which represent the contribution of trajectory errors to the objective function value. In related technologies, the penalty weights are often fixed, and the objective function is constructed based on these fixed values ​​combined with a quadratic penalty mechanism.

[0004] However, in real-world complex driving scenarios, the vehicle's operating environment is constantly changing. In the aforementioned related technologies, the penalty weights are fixed, and the penalty mechanism in the objective function is of a single form. This results in the constructed objective function being unable to adapt to the dynamically changing vehicle driving state, leading to unreasonable planned vehicle driving trajectories and even causing safety hazards. Summary of the Invention

[0005] This specification provides a vehicle trajectory planning method, apparatus, and electronic device to adapt to dynamically changing vehicle driving states and improve the safety and rationality of the planned vehicle trajectory.

[0006] This specification provides an embodiment of a vehicle trajectory planning method, including: The configuration mode for penalty weighting is determined based on vehicle driving data; Based on the configuration mode, determine the value of the penalty weight; Based on the value of the penalty weight, an objective function is constructed. The objective function is used to measure the quality of the driving trajectory through multiple piecewise penalty functions with multiple optimization dimensions. The multiple function branches of the piecewise penalty function correspond to multiple penalty methods. The objective function is optimized and solved using a solver to obtain the driving trajectory.

[0007] In some embodiments, the configuration mode for determining the penalty weight includes: Identify vehicle driving scenarios based on vehicle driving data; Select the configuration mode that matches the vehicle driving scenario.

[0008] In some embodiments, identifying the vehicle driving scenario includes: When the vehicle driving data meets the preset rules, the scene corresponding to the preset rules is obtained as the vehicle driving scene; When the vehicle driving data does not meet the preset rules, the vehicle driving data is predicted by a machine learning model to obtain the vehicle driving scenario output by the machine learning model.

[0009] In some embodiments, selecting a configuration mode that matches the vehicle driving scenario includes: When the confidence level of the vehicle driving scenario is less than or equal to a preset confidence threshold, the configuration mode is determined to be a downgraded mode. The penalty weight is set to a preset calibration value.

[0010] In some embodiments, the method further includes: Calculate the penalty method switching threshold based on the vehicle driving data; The objective function to be constructed includes: Based on the values ​​of the penalty weights and the threshold for switching the penalty method, a target function is constructed.

[0011] In some embodiments, the piecewise penalty function includes a first function branch corresponding to a quadratic penalty method and a second function branch corresponding to a linear penalty method; The first function branch includes a quadratic penalty weight; The second function branch includes linear penalty weights.

[0012] In some embodiments, the piecewise penalty function is ; Let i represent the piecewise penalty function for the i-th optimization dimension. This is the first function branch. This is the second function branch. This represents the threshold for switching the penalty method in the i-th optimization dimension. This represents the quadratic penalty weight for the i-th optimization dimension. This represents the linear penalty weight for the i-th optimization dimension. This represents the trajectory error of the driving trajectory in the i-th optimization dimension.

[0013] In some embodiments, the configuration mode is a learning mode; The process of determining the penalty weight value based on the configuration mode, constructing an objective function based on the penalty weight value, optimizing and solving the objective function to obtain the driving trajectory includes: Iteratively execute the following steps until the preset conditions are met: Using the Bayesian optimization algorithm, candidate values ​​for the penalty weight are selected; Construct the objective function based on the candidate values ​​of the penalty weights; The objective function is optimized and solved using a solver to obtain the solution result; Based on the solution results, calculate the performance index of the candidate values; Update the optimal value of the penalty weight based on the performance metrics of the candidate values; After the iteration is complete, output the driving trajectory corresponding to the optimal value of the penalty weight.

[0014] In some embodiments, the method further includes: Calculate the performance metric for the candidate values ​​based on at least one of the following: The convergence speed of the solver, the feasibility of the driving trajectory, and the quality of the driving trajectory.

[0015] In some embodiments, the multiple optimization dimensions include a safety dimension, the candidate values ​​include the penalty weight values ​​under the safety dimension, and the solution result is a solution failure; the method further includes: By increasing the penalty weight values ​​under the security dimension, we obtain the adjusted values; Based on the adjusted candidate values ​​of the penalty weights, the objective function construction steps and optimization solution steps are repeated.

[0016] In some embodiments, the method further includes: If the number of failures reaches the preset number of failures, the configuration mode will be changed from learning mode to degraded mode. Based on the preset values ​​of the penalty weights, the objective function construction steps and optimization solution steps are repeated.

[0017] In some embodiments, the configuration mode is a normal mode or an extreme mode; The determination of the penalty weight value includes: Based on vehicle driving data, the penalty weight values ​​are matched in the parameter set. The parameter set includes multiple subsets, and the subsets include the values ​​of the penalty weights.

[0018] In some embodiments, the configuration mode is a normal mode; The piecewise penalty function includes a first function branch corresponding to a quadratic penalty method and a second function branch corresponding to a linear penalty method. The first function branch includes quadratic penalty weights, and the second function branch includes linear penalty weights. The parameter set includes a first parameter set, which includes multiple first subsets. The first subsets include the values ​​of quadratic penalty weights and linear penalty weights. In the first subsets, the value of the quadratic penalty weight is greater than the value of the linear penalty weight.

[0019] In some embodiments, the configuration mode is an extreme mode; The piecewise penalty function includes a first function branch corresponding to a quadratic penalty method and a second function branch corresponding to a linear penalty method. The first function branch includes quadratic penalty weights, and the second function branch includes linear penalty weights. The parameter set includes a second parameter set, which includes multiple second subsets. The second subsets include the values ​​of quadratic penalty weights and linear penalty weights. In the second subsets, the value of the quadratic penalty weight is less than the value of the linear penalty weight.

[0020] This specification also provides a vehicle trajectory planning device, comprising: The first determining unit is used to determine the configuration mode of the penalty weight based on the vehicle driving data; The second determining unit is used to determine the value of the penalty weight according to the configuration mode; The construction unit is used to construct an objective function based on the value of the penalty weight. The objective function is used to measure the quality of the driving trajectory through multiple piecewise penalty functions with multiple optimization dimensions. The multiple function branches of the piecewise penalty function correspond to multiple penalty methods. The solution unit is used to optimize and solve the objective function using a solver to obtain the driving trajectory.

[0021] This specification also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described vehicle trajectory planning method.

[0022] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described vehicle trajectory planning method.

[0023] This specification also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described vehicle trajectory planning method.

[0024] The vehicle trajectory planning method described in this specification can determine the configuration mode of penalty weights based on vehicle driving data; determine the value of the penalty weights based on the configuration mode; construct an objective function based on the value of the penalty weights, the objective function being used to measure the quality of the driving trajectory through multiple piecewise penalty functions with various optimization dimensions, and the multiple function branches of the piecewise penalty functions corresponding to various penalty methods; and optimize the objective function using a solver to obtain the driving trajectory. This specification embodiment can determine the configuration mode of penalty weights based on vehicle driving data, and determine the value of the penalty weights based on the configuration mode. This achieves dynamic adaptation of the penalty weight value to the vehicle driving state. Furthermore, the objective function can evaluate the quality of the driving trajectory through multiple piecewise penalty functions with various optimization dimensions, and the multiple function branches of the piecewise penalty functions correspond to various penalty methods, facilitating the selection of the appropriate function branch based on the vehicle driving data. This achieves dynamic adaptation of the penalty method to the vehicle driving state. The embodiments in this specification effectively avoid the problem of unreasonable trajectory planning caused by fixed penalty weight values ​​and a single penalty method by dynamically adapting the penalty weight values ​​and penalty methods to the environment. This significantly improves the adaptability and operational reliability of the vehicle's autonomous driving system in complex road environments. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. The drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a flowchart illustrating the vehicle trajectory planning method in the embodiments of this specification; Figure 2 This is a flowchart illustrating the vehicle trajectory planning method in the embodiments of this specification; Figure 3 This is a flowchart illustrating the driving trajectory planning method in the learning mode as described in the embodiments of this specification. Figure 4 This is a functional structure diagram of the vehicle trajectory planning method in the embodiments of this specification. Detailed Implementation

[0027] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. The specific embodiments described herein are only used to explain this disclosure, and not to limit this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure are within the scope of protection of this disclosure. In addition, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0028] The aforementioned technologies have the following problems.

[0029] (1) Insufficient robustness of numerical optimization. The quadratic penalty mechanism amplifies errors quadratically. When dealing with outlier scenarios (such as sudden obstacle intrusion or a sharp drop in road friction coefficient), the fixed penalty intensity is prone to mismatch with the scenario requirements, leading to a surge in the condition number of the objective function and causing numerical oscillations. This can cause abnormal convergence of the optimization solver, a decrease in the feasibility of the solution, or even solution failure, seriously affecting the safety of vehicle decision-making.

[0030] (2) Poor adaptability to complex scenarios. The single value of the penalty weight is difficult to take into account the optimization needs of various scenarios, resulting in the planned trajectory being "overly conservative" or "overly aggressive" in some scenarios.

[0031] (3) Lack of online self-learning and dynamic adjustment capabilities. The value of the penalty weights mostly depends on offline calibration, and the optimal value is determined through a large number of real vehicles or simulation scenarios. It is impossible to adjust online according to real-time scenario changes. When encountering extreme scenarios that are not covered (such as icy roads, construction road occupancy + sudden pedestrians), fixed values ​​can easily lead to a significant decrease in planning performance and even cause safety hazards.

[0032] (4) The penalty function has a single form and the error handling mechanism is imperfect. Traditional planning algorithms mostly use pure quadratic penalty functions, which can ensure the smoothness of optimization in normal scenarios. However, in outlier or extreme error scenarios, higher-order terms will excessively amplify the impact of errors, causing the optimization direction to deviate from the actual needs. There is a lack of flexible penalty form switching mechanism, making it impossible to achieve a dynamic balance between "smoothness" and "anti-interference".

[0033] This specification provides a vehicle trajectory planning method. The method can be applied to electronic devices, including autonomous driving domain controllers. Please refer to [link to documentation]. Figure 1 The method may include the following steps.

[0034] Step 11: Determine the configuration mode of the penalty weight based on the vehicle driving data; Step 12: Determine the value of the penalty weight according to the configuration mode; Step 13: Based on the value of the penalty weight, construct the objective function. The objective function is used to measure the quality of the driving trajectory through multiple piecewise penalty functions with multiple optimization dimensions. The multiple function branches of the piecewise penalty function correspond to multiple penalty methods. Step 13: Use the solver to optimize the objective function and obtain the driving trajectory.

[0035] This embodiment of the specification can determine the configuration mode of the penalty weight based on vehicle driving data, and determine the value of the penalty weight based on the configuration mode. This achieves dynamic adaptation of the penalty weight value to the vehicle driving state. Furthermore, the objective function can evaluate the quality of the driving trajectory through multiple segmented penalty functions with various optimization dimensions. The multiple function branches of these segmented penalty functions correspond to various penalty methods, facilitating the selection of the appropriate function branch based on vehicle driving data. This achieves dynamic adaptation of the penalty method to the vehicle driving state. By dynamically adapting the penalty weight value and penalty method to the environment, this embodiment of the specification effectively avoids the problem of unreasonable trajectory planning caused by fixed penalty weight values ​​and a single penalty method, significantly improving the adaptability and operational reliability of the vehicle's autonomous driving system in complex road environments.

[0036] In some embodiments, vehicle driving data is used to represent the state of the vehicle while it is in motion.

[0037] The vehicle driving data may include obstacle data, road data, vehicle driving status data, traffic rule data, and vehicle interaction data. Obstacle data may include the distance between the vehicle and the obstacle, the type of obstacle, the vehicle's speed relative to the obstacle, and the vehicle's movement trend relative to the obstacle. The vehicle's movement trend relative to the obstacle is selected from "moving away" and "approaching." Road data may include road structure data and road surface condition data. Road structure data may include lane markings, curvature, and slope. Road surface condition data may include friction coefficient, water accumulation markings, ice accumulation markings, and snow accumulation markings. Vehicle driving status data may include vehicle speed, acceleration, and yaw rate. Traffic rule data may include speed limits, right-of-way rules, and construction zones. Vehicle interaction data may include the positions of surrounding vehicles and their driving intentions.

[0038] Obstacle and road data can be collected through the vehicle's environmental perception components. These components may include lidar, cameras, millimeter-wave radar, etc. Vehicle driving status data can be collected through vehicle status sensors. Traffic rule data and vehicle interaction data can be obtained through V2X (Vehicle-to-Everything) and map data.

[0039] In some embodiments, feature engineering can be performed on vehicle driving data to obtain driving feature data. Driving feature data may include multiple sub-feature data. These sub-feature data may include the distance between the vehicle and the nearest obstacle, the road surface friction coefficient, the vehicle speed, the number of surrounding interacting vehicles, and road curvature. The number of surrounding interacting vehicles reflects the density of multi-vehicle interactions, and road curvature reflects the path complexity. For example, the driving feature data may include a feature vector. The feature vector F = [d...] obst , μ, v self , n veh , k road ]. d obst This represents the distance between the vehicle and the nearest obstacle, which may include lateral distance and / or longitudinal distance, μ represents the road surface friction coefficient, and v self n represents the vehicle speed. veh k represents the number of vehicles interacting with the surrounding area. road Indicates the curvature of the road.

[0040] Optionally, the driving characteristic data can also be smoothed and normalized. For example, the d... obst v self Sub-feature data are smoothed. Smoothing removes measurement noise. For example, each sub-feature data can be mapped to the [0,1] interval. Normalization ensures consistency of classifier input.

[0041] In some embodiments, the configuration mode for penalty weights can be determined directly based on vehicle driving data, or it can be determined based on driving characteristic data. Vehicle driving scenarios can be identified based on vehicle driving data or driving characteristic data, and a configuration mode matching the vehicle driving scenario can be selected.

[0042] Vehicle driving scenarios can include conventional driving scenarios, extreme driving scenarios, and multi-vehicle interactive driving scenarios. Conventional driving scenarios represent low-risk environments with no obstacles, good road surfaces, and smooth traffic. In conventional driving scenarios, a high degree of trajectory smoothness is required for a more comfortable ride. Extreme driving scenarios represent high-risk environments such as vehicles encountering close-range obstacles, low-friction road surfaces (such as ice, snow, or standing water), and emergency avoidance. In extreme driving scenarios, a high degree of trajectory safety is required. Multi-vehicle interactive driving scenarios represent environments where vehicles are in congested following, lane-changing maneuvers, or crossing intersections, where multiple vehicles interact and influence each other. In multi-vehicle interactive driving scenarios, it is necessary to dynamically adjust penalty weights to balance traffic efficiency and cooperation.

[0043] A hybrid approach combining rules and machine learning can be used to identify vehicle driving scenarios. Specifically, one or more preset rules can be pre-configured, each corresponding to a different driving scenario. The system can determine whether vehicle driving data or driving feature data satisfies the one or more preset rules. When the vehicle driving data or driving feature data satisfies the preset rules, the scenario corresponding to those rules is obtained as the vehicle driving scenario. When the vehicle driving data or driving feature data does not satisfy any preset rules, a machine learning model can be used to predict the vehicle driving scenario output by the machine learning model.

[0044] Therefore, preset rules can be prioritized for rapid identification of typical vehicle driving scenarios, ensuring real-time and deterministic response in critical scenarios. When preset rules cannot cover all scenarios, machine learning models can be used for supplementary identification, thereby improving the coverage and accuracy of scenario classification while balancing recognition efficiency and generalization ability.

[0045] For example, preset rules correspond to extreme driving scenarios. Preset rules may include: d obst The value μ is less than or equal to a first preset value, and μ is less than or equal to a second preset value. For example, the first preset value could be 5m, and the second preset value could be 0.3. When the driving feature data meets the preset rules, the vehicle driving scenario can be determined to be an extreme driving scenario. For example, machine learning models include random forest models. Driving feature data can be input into a machine learning model to obtain the vehicle driving scenario output by the model. The vehicle driving scenario is selected from normal driving scenarios, extreme driving scenarios, and multi-vehicle interaction driving scenarios.

[0046] Configuration modes include Normal Mode, Extreme Mode, Learning Mode, and Fallback Mode. There is a correspondence between vehicle driving scenarios and configuration modes. The configuration mode with the appropriate penalty weight can be determined based on this correspondence.

[0047] For example, there is a correspondence between normal driving scenarios and normal modes, between extreme driving scenarios and extreme modes, and between multi-vehicle interactive driving scenarios and learning modes. The configuration mode can be determined as normal mode, extreme mode, or learning mode based on whether it is a normal driving scenario, an extreme driving scenario, or a multi-vehicle interactive driving scenario.

[0048] Optionally, when identifying vehicle driving scenarios using a machine learning model, the model can output the vehicle driving scenario and its confidence level. The confidence level represents the degree of credibility of the vehicle driving scenario output by the machine learning model. The confidence level is positively correlated with the degree of credibility. Therefore, a confidence threshold can be preset; the confidence level of the vehicle driving scenario output by the machine learning model can be compared with the confidence threshold. If the confidence level is greater than or equal to the confidence threshold, the corresponding configuration mode can be determined based on the vehicle driving scenario output by the machine learning model; if the confidence level is less than the confidence threshold, the configuration mode can be determined to be a degraded mode. The confidence threshold can be, for example, 0.7. Thus, when the confidence level of the machine learning model's prediction result is low, it can automatically switch to a degraded mode, avoiding unreasonable driving trajectory planning due to misidentification of driving scenarios, thereby improving the safety and robustness of the autonomous driving system in complex environments.

[0049] In some embodiments, the objective function is used to measure the quality of a driving trajectory through multiple piecewise penalty functions of various optimization dimensions. The objective function may include multiple piecewise penalty functions. Each piecewise penalty function corresponds to an optimization dimension. The multiple piecewise penalty functions include a piecewise penalty function corresponding to the safety dimension (hereinafter referred to as the first piecewise penalty function), a piecewise penalty function corresponding to the smoothness dimension (hereinafter referred to as the second piecewise penalty function), a piecewise penalty function corresponding to the efficiency dimension (hereinafter referred to as the third piecewise penalty function), and a piecewise penalty function corresponding to the cooperation dimension (hereinafter referred to as the fourth piecewise penalty function).

[0050] The objective function can be obtained by weighted summing of the piecewise penalty functions for each optimization dimension. For example, the objective function J... tatal =J safe +λ1×J smooth +λ2×J eff +λ3×J coop J safeLet J represent the first segment penalty function. smooth J represents the second piecewise penalty function. eff J represents the third piecewise penalty function. coop This represents the fourth piecewise penalty function. λ1, λ2, and λ3 are the weights of the optimization dimensions. λ1 is the weight of the smoothness dimension, λ2 is the weight of the efficiency dimension, and λ3 is the weight of the synergy dimension. The optimization dimension weights are used to weighted sum the piecewise penalty functions of each optimization dimension, representing the relative importance of the piecewise penalty function of each optimization dimension in the objective function. The optimization dimension weights can be fixed, or they can be dynamically adjusted according to the driving scenario. For example, the value of λ1 is greater in normal driving scenarios than in extreme driving scenarios, while the value of λ2 is less in normal driving scenarios than in extreme driving scenarios. Thus, in extreme driving scenarios, the value of λ1 decreases, and the value of λ2 increases.

[0051] Different piecewise penalty functions correspond to different optimization dimensions, penalizing vehicle trajectories from different perspectives. The independent variable of a piecewise penalty function can be the trajectory error, and its value can be understood as a penalty for that error. The trajectory error is the deviation of the vehicle's trajectory within the optimization dimension, and can include safety error, smoothing error, efficiency error, and cooperative error. The objective function's value can be understood as the total penalty for the vehicle's trajectory. For example, the first piecewise penalty function penalizes the vehicle's trajectory from a safety perspective. The trajectory error of the first piecewise penalty function is the safety error, reflecting the safety of the trajectory. The safety error can be the distance between the vehicle and obstacles. The first piecewise penalty function measures the risk of collision between the vehicle and obstacles, encouraging vehicles to maintain a safe distance. The second piecewise penalty function penalizes the vehicle's trajectory from a smoothing perspective. The trajectory error of the second piecewise penalty function is the smoothing error, reflecting the smoothness of the trajectory. The smoothing error can be calculated based on the trajectory curvature change rate error, vehicle acceleration error, etc. For example, the smoothing error can be obtained by weighted fusion of the trajectory curvature change rate error and the vehicle acceleration error. The trajectory curvature change rate error can be the deviation between the actual curvature change rate and the expected curvature change rate. The vehicle acceleration error can be the deviation between the trajectory acceleration and the expected acceleration. The second-segment penalty function measures the curvature change rate of the vehicle's trajectory, encouraging a smooth trajectory for a more comfortable ride. The third-segment penalty function penalizes the vehicle's trajectory from an efficiency perspective. The trajectory error of the third-segment penalty function is the efficiency error, reflecting vehicle driving efficiency; it can be the deviation between the actual vehicle speed and the expected speed. The third-segment penalty function measures the deviation between the vehicle's trajectory speed and the expected speed, encouraging efficient vehicle movement. The fourth-segment penalty function penalizes the vehicle's trajectory from a cooperation perspective. The trajectory error of the fourth-segment penalty function can be the cooperation error, reflecting the degree of multi-vehicle interaction and cooperation. The cooperation error can be calculated based on distance to surrounding vehicles, speed difference, and interaction intention deviation. For example, the cooperation error can be obtained by weighted fusion of distance to surrounding vehicles, speed difference, and interaction intention deviation. Interaction intention deviation can be the degree of difference between the trajectory and the expected driving intention of surrounding vehicles. The fourth segment penalty function can measure the degree of coordination between the vehicle's driving trajectory and the driving intentions of other vehicles in multi-vehicle interaction scenarios, and encourage reasonable distance maintenance and courteous behavior.

[0052] In the objective function, each piecewise penalty function (e.g., the first piecewise penalty function, the second piecewise penalty function, the third piecewise penalty function, the fourth piecewise penalty function, etc.) can include multiple function branches. Each function branch can be understood as a penalty branch. The piecewise penalty function includes a first function branch corresponding to a quadratic penalty method and a second function branch corresponding to a linear penalty method. The first function branch includes a quadratic function and the penalty weight corresponding to the quadratic function (hereinafter referred to as quadratic penalty weight). The second function branch includes a linear function and the penalty weight corresponding to the linear function (hereinafter referred to as linear penalty weight). Optionally, the second function branch may also include quadratic penalty weights, but not the quadratic function. Thus, through the piecewise penalty function, the dynamic switching and fusion of quadratic and linear penalties can be achieved.

[0053] The piecewise penalty function implements a piecewise fusion mechanism of quadratic and linear penalties, which can balance the smoothness of normal scenarios and the robustness to interference in extreme scenarios. The piecewise penalty function avoids numerical oscillations in outlier scenarios caused by pure quadratic penalties, can ensure the efficient convergence of the solver, improve the feasibility of solutions, and thus significantly improve the robustness of numerical optimization.

[0054] For example, the piecewise penalty function can be ; Let i represent the piecewise penalty function for the i-th optimization dimension. This is the first function branch. This is the second function branch. This represents the threshold for switching the penalty method in the i-th optimization dimension. This represents the quadratic penalty weight for the i-th optimization dimension. This represents the linear penalty weight for the i-th optimization dimension. This represents the trajectory error of the driving trajectory in the i-th optimization dimension. 1 ≤ i ≤ 4, where i is a positive integer. The first optimization dimension can be a safety dimension, the second a smoothness dimension, the third an efficiency dimension, and the fourth a cooperative dimension. Additionally, for the piecewise penalty function... When i is 1, 2, 3, or 4, i represents safe, smooth, eff, and coop, respectively, representing the piecewise penalty function. Correspondingly, these can be the first segment penalty function, the second segment penalty function, the third segment penalty function, and the fourth segment penalty function, respectively. Correspondingly, ω safe1 ω smooth1 ω eff1 ω coop1 , Correspondingly, ω safe2 ω smooth2 ω eff2ω coop2 , Correspondingly, e safe e smooth e eff e coop , Correspondingly, θ safe θ smooth θ eff θ coop .

[0055] In some embodiments, penalty weights may specifically include weights in the objective function used to represent the contribution of the penalty method to the objective function value (penalty value). The objective function may include multiple penalty weights, which can form a penalty weight set for the objective function. The penalty weight set may include multiple sets of penalty weights. Each set of penalty weights corresponds to an optimization dimension, including the penalty weights in the piecewise penalty function corresponding to that optimization dimension. The piecewise penalty function may include quadratic penalty methods and linear penalty methods. Therefore, in the penalty weight set, each set of penalty weights may include multiple penalty weights, for example, quadratic penalty weights and linear penalty weights. Quadratic penalty weights are weights for quadratic penalty methods, and linear penalty weights are weights for linear penalty methods. Quadratic penalty methods are represented by quadratic functions, and linear penalty methods are represented by linear functions.

[0056] For example, the penalty weight set of the objective function may include W=[ω safe1 ω safe2 ω smooth1 ω smooth2 ω eff1 ω eff2 ω coop1 ω coop2 ]. Wherein, ω safe1 ω safe2 ω represents a set of penalty weights, corresponding to the security dimension. safe1 For quadratic penalty weights, ω safe2 ω represents the linear penalty weight. smooth1 ω smooth2 ω represents a set of penalty weights, corresponding to the smoothing dimension. smooth1 For quadratic penalty weights, ω smooth2 ω represents the linear penalty weight. eff1 ω eff2 ω represents a set of penalty weights, corresponding to the efficiency dimension. eff1 For quadratic penalty weights, ω eff2 ω represents the linear penalty weight. coop1 ω coop2 ω represents a set of penalty weights, corresponding to the collaborative dimension. coop1 For quadratic penalty weights, ω coop2The weights are linear penalty weights.

[0057] In some embodiments, the objective function may further include multiple penalty mode switching thresholds, each corresponding to a piecewise penalty function, thereby corresponding to an optimization dimension. The penalty mode switching thresholds are used to distinguish multiple function branches of the piecewise penalty function, thus enabling the differentiation of various penalty modes within the piecewise penalty function.

[0058] In the objective function, the threshold for switching the penalty method in each segment of the penalty function can be fixed, or it can be adaptively adjusted based on the vehicle driving data. Therefore, the threshold for switching each penalty method can be calculated based on the vehicle driving data; and the objective function can be constructed based on the values ​​of each penalty weight and the threshold for switching each penalty method.

[0059] The first segment penalty function corresponds to the safety dimension. It can be based on d. obst Determine the penalty mode switching threshold for the first segment penalty function. This penalty mode switching threshold can be related to d. obst It shows a negative correlation. Therefore, d obst The smaller the value of μ, the smaller the threshold for switching the penalty method, and the earlier the linear penalty is switched. The second piecewise penalty function corresponds to the smoothness dimension. The penalty method switching threshold of the second piecewise penalty function can be determined based on μ. This penalty method switching threshold can be positively correlated with μ. Therefore, the larger μ is, the larger the penalty method switching threshold is, prioritizing driving stability. The third piecewise penalty function corresponds to the efficiency dimension. The threshold for switching the penalty method can be determined based on v. self Determine the penalty mode switching threshold for the third segment penalty function. This penalty mode switching threshold can be related to v. self They are positively correlated. Therefore, v self The larger the value, the higher the threshold for switching the penalty method, prioritizing driving efficiency. The fourth segment penalty function corresponds to the collaborative dimension. This can be determined based on n. veh Determine the penalty mode switching threshold for the fourth segment penalty function. This penalty mode switching threshold can be related to n. veh They are positively correlated. Therefore, n veh The larger the value, the higher the threshold for switching the penalty method, balancing smoothness and collaborative safety. In practical applications, vehicle driving data can be substituted into the switching threshold function to calculate the penalty method switching threshold. For example, d can be... obst Substituting μ into the first switching threshold function, the penalty mode switching threshold of the first segmented penalty function is calculated; μ can be substituted into the second switching threshold function to calculate the penalty mode switching threshold of the second segmented penalty function; v can be... self Substituting into the second switching threshold function, the penalty mode switching threshold of the third segmented penalty function is calculated; n can be used as the threshold for switching. vehSubstituting the values ​​into the second switching threshold function, the penalty mode switching threshold of the fourth segment penalty function is calculated. Of course, other methods can also be used to determine the penalty mode switching threshold of each segment penalty function; this specification does not limit this approach.

[0060] In some embodiments, the objective function may further include one or more optimization dimension weights, such as weights for safety, smoothness, efficiency, and collaboration. In the objective function, the weights of each optimization dimension can be fixed or adaptively adjusted. Therefore, vehicle driving scenarios can be identified based on vehicle driving data or driving characteristic data, and the values ​​of the optimization dimension weights can be determined based on the vehicle driving scenarios. The objective function can be constructed based on the values ​​of each penalty weight, the threshold for switching each penalty method, and the values ​​of each optimization dimension weight.

[0061] Multiple value sets can be provided. Each value set corresponds to a vehicle driving scenario and includes one value for the weights of multiple optimization dimensions, such as values ​​for safety, smoothness, efficiency, and collaboration. The same optimization dimension weight may have different values ​​in different value sets. A target value set can be selected from these multiple value sets based on the vehicle driving scenario; the values ​​of each optimization dimension weight can be obtained from the target value set.

[0062] In some embodiments, please refer to Figure 2 The configuration mode is used to represent the strategy for determining the values ​​of each penalty weight in the penalty weight set. The strategy for determining the values ​​of each penalty weight in the penalty weight set can also differ under different configuration modes.

[0063] In learning mode, the optimal penalty weights can be searched online using a Bayesian optimization algorithm, yielding the corresponding driving trajectory. This online search for optimal penalty parameter combinations eliminates the need for manual intervention, allowing for adaptive adjustment of penalty parameters and reducing the cost of extensive offline calibration. The integration of scene perception and Bayesian optimization enables real-time matching of penalty parameters with scene features, overcoming the limitations of fixed penalty parameters in complex scenarios. This allows for the generation of optimal trajectories even in situations involving sudden obstacles, low-adhesion surfaces, and multi-vehicle interactions. In learning mode, the optimal penalty weights can be searched online using a Bayesian optimization algorithm, yielding the corresponding driving trajectory. In normal or extreme modes, the optimal penalty weights can be selected from a parameter set, yielding the corresponding driving trajectory. In degraded mode, preset calibration values ​​can be obtained, along with the corresponding driving trajectory. When scene recognition fails or the solution fails, robust parameters are pre-calibrated to ensure uninterrupted vehicle operation, thereby improving the reliability and safety of autonomous driving.

[0064] In some embodiments, the configuration mode is learning mode.

[0065] Please see Figure 3 The following steps can be executed iteratively until the preset conditions are met: Using the Bayesian optimization algorithm, candidate values ​​for the penalty weight are selected; based on the candidate values ​​for the penalty weight, an objective function is constructed; the objective function is optimized and solved using a solver to obtain the solution result; based on the solution result, the performance index of the candidate values ​​is calculated; based on the performance index of the candidate values, the optimal value of the penalty weight is updated.

[0066] After the iteration is complete, output the driving trajectory corresponding to the optimal value of the penalty weight.

[0067] The following describes the process of selecting candidate values ​​for penalty weights using the Bayesian optimization algorithm.

[0068] Candidate values ​​for penalty weights can be selected using a parametric evaluation function. This function measures the quality of the penalty weight values. Considering the difficulty in determining the expression for the parametric evaluation function, a surrogate model can be used. This surrogate model simulates the parametric evaluation function. The surrogate model can be updated iteratively to improve its fitting accuracy to the parametric evaluation function. The surrogate model can include a Gaussian process regression (GPR) model. This model can be characterized by a mean function and a kernel function, such as the Matrn kernel function or the quadratic exponential kernel function. The surrogate model can output a predicted mean and a predicted standard deviation for any candidate penalty weight value. The predicted mean represents the estimated value of the parametric evaluation function at that value, and the predicted standard deviation represents the uncertainty of that estimate.

[0069] A proxy model can be trained based on the current sample set. The current sample set can include one or more subsets. Each subset can include a value for each penalty weight in the penalty weight set, and a corresponding evaluation value. The same penalty weight can have the same or different values ​​in different subsets. The following example illustrates the training process of the proxy model.

[0070] You can retrieve the current sample set. The current sample set can include (W1, y1), (W2, y2), ..., (W... m , y m W1, W2, W m Denotes a subset of samples, y1, y2, y3, y4, y5, y6 m This represents the evaluation value. m represents the number of subsets in the current sample set.

[0071] Gaussian process f(W)~ It can be used as a proxy model to fit the penalty weight values ​​and the evaluation values. Represents the mean function, Let f(W) represent the covariance function (kernel function). Let f(W) represent the parameter evaluation function. s represents the dimension of the penalty weight set W, that is, the number of penalty weights in the objective function. Let represent the value of the d-th penalty weight in the p-th subset. For signal variance, Let represent the length scale of the d-th penalty weight. Considering observation noise, assume the observation value... Interference from independent and identically distributed Gaussian noise, , At this point, the two observations and The covariance between them is , Represents the Kronecker function (when hour =1, otherwise (0).

[0072] Hyperparameters to be trained .

[0073] Based on the current sample set, calculate kernel matrix . p,q=1,…,m. After considering observation noise, the covariance matrix of the observed data is: . yes The identity matrix.

[0074] make Given hyperparameters Below, observation data The marginal likelihood of is in Gaussian form, and its logarithmic expression is: . Used to measure how well the surrogate model fits the observed data. This is a penalty term for the complexity of the surrogate model, used to prevent overfitting. This is a constant normalization term. The hyperparameters can be estimated by maximizing the logarithmic marginal likelihood. Specifically, numerical optimization methods (such as gradient ascent, L-BFGS, etc.) can be used to maximize this. The function can be calculated with respect to... The gradient of each hyperparameter is calculated. These gradients can be analytically obtained using the partial derivatives of the kernel function with respect to the hyperparameters. The optimization process is iterative until the logarithmic marginal likelihood converges or the preset maximum number of iterations is reached. The final optimal hyperparameters can be denoted as... .

[0075] The value space of the penalty weights can be determined. For example, the range of quadratic penalty weights is [0.1, 5.0], and the range of linear penalty weights is [1.0, 10.0]. The constraints on the penalty weights can be determined. For example, the constraints include: the value of the linear penalty weight is greater than the value of the quadratic penalty weight. Within the above value space, one or more possible values ​​of each penalty weight in the penalty weight set that satisfy the above constraints can be searched to obtain one or more subsets. Each subset includes one possible value of each penalty weight in the penalty weight set. For example, one or more subsets can be obtained by uniform sampling. Then, for each subset, the prediction mean and prediction variance can be obtained through a surrogate model. For example, for any subset... f( )| ~ . . . Indicates proxy model The predicted mean, Indicates proxy model The prediction variance. .

[0076] Based on the predicted mean and predicted variance, a subset can be selected as the target subset from one or more subsets using a collection function. Specifically, each subset can be substituted into the collection function to obtain the collection function value for that subset; the subset with the largest collection function value can be selected as the target subset from the one or more subsets. Values ​​in the target subset can be used as candidate values ​​for penalty weights. The collection function can include an Expected Improvement (EI) function. For example, if... Then there is ;if Then there is . and These represent the proxy model targeting... The output shows the predicted mean and predicted standard deviation. This represents the largest evaluation value among all sub-sample sets in the current sample set. and These are the cumulative distribution function and probability density function of the standard normal distribution, respectively.

[0077] The following describes the process of constructing the objective function.

[0078] The target subset can include multiple candidate values ​​for penalty weights. The current driving trajectory can be obtained; based on the current driving trajectory, the target subset, the switching thresholds for each penalty method, and the values ​​of the weights for each optimization dimension, the objective function J can be constructed. tatal =J safe +λ1×J smooth +λ2×J eff +λ3×J coop The current driving trajectory can include the vehicle's position, speed, acceleration, steering angle sequence, etc.

[0079] The current driving trajectory can be obtained. The current driving trajectory can be discretized into N sub-trajectories over N time steps. N can be, for example, 30. The current driving trajectory can include N sub-trajectories over N time steps, and each sub-trajectory over N time steps can include position, velocity, acceleration, steering angle sequence, etc. For example, the current driving trajectory... . Indicates location, Indicates speed, Indicates acceleration. Indicates the steering angle.

[0080] Constraints can be obtained. Constraints include vehicle dynamics constraints, safety constraints, and road constraints. For example, vehicle dynamics constraints are used to constrain maximum steering angle and maximum acceleration, safety constraints are used to constrain minimum distances to obstacles, and road constraints are used to constrain lane boundaries. For instance, dynamics constraints may include: , , , Safety constraints may include: the distance between the vehicle and the obstacle is not less than the minimum safe distance. Road constraints may include: vehicle position. It must not exceed the lane lines. The above constraints can be uniformly written as inequalities. Sum of equations .

[0081] Objective function can be constructed . Let represent the first, second, third, and fourth segment penalty functions at the k-th time step, respectively. For example, Specifically, for each time step It is possible to extract the sub-trajectory of a given time step from the current driving trajectory. The sub-trajectory of that time step can include... Based on the sub-trajectories at a given time step, the trajectory error for each optimization dimension can be calculated. The trajectory error for each optimization dimension can include... Based on the trajectory error of each optimization dimension, combined with the target subset and the switching threshold of each penalty method, a piecewise fusion penalty function for each optimization dimension can be constructed. The piecewise fusion penalty function for each optimization dimension can include... The objective function for a given time step can be obtained by combining the piecewise fusion penalty function for each optimization dimension with the values ​​of the weights for each optimization dimension. The objective functions at each time step can be summed to obtain... .

[0082] The following describes the process of optimizing the objective function using a solver. The solver may include a fast quadratic programming solver (such as OSQP), a nonlinear programming solver (such as IPOPT), etc. The optimization process may involve multiple iterations.

[0083] The optimization iteration can be terminated if any of the following conditions are met: optimality condition, maximum number of iterations, or timeout. Optimality condition may include: the norm of the objective function gradient. And all constraint violations (generally Timeout can include: exceeding a preset time since the start of the solution. (For example ).

[0084] The search direction can be calculated. Specifically, the sequential quadratic programming (SQP) method can be used to construct the quadratic programming sub-functions: , .in, Let be the Hessian matrix of the Lagrange function. This is the gradient of the objective function at the current point. Let be the Jacobian matrix of the constraints. The search direction can be obtained by solving the above quadratic programming sub-function. .

[0085] You can search online to determine the step size. .from Initially, a backtracking method was used to narrow down the possibilities. This continues until a search stopping condition is met (e.g., the Armijo sufficient descent condition). The Armijo sufficient descent condition is... .in, . For the sub-trajectory of the new time step. It can be based on Update the current driving trajectory x to obtain the solution. For example, you can... Add the current driving trajectory x to obtain the solution result. The solution result can include the updated current driving trajectory. Therefore, the updated current driving trajectory can be obtained. The updated current driving trajectory can be represented as... The updated current driving trajectory can be understood as adding one or more new sub-trajectories to the current driving trajectory, such as adding a sub-trajectory at the N+1th time step. .

[0086] It can also verify whether the solution results meet the constraints.

[0087] The performance metrics for candidate values ​​(target subsets) can be calculated based on at least one of the following: solver convergence speed, trajectory feasibility, and trajectory quality. For example, this can be achieved using the formula... Obj ( )= α Conv + β Feas + γ Qual Calculate the performance metrics for candidate values. Indicates the target subset. Obj ( The ) represents the performance index of the target subset. α, β, and γ are weighting coefficients, with β>α>γ, prioritizing feasibility. Conv represents the convergence speed of the solver, which can be normalized to the interval [0,1]. A larger Conv indicates faster convergence. Feas represents the feasibility of the solution; Feas=0 indicates the solution is infeasible (the solution does not meet the constraints), and Feas=1 indicates the solution is feasible (the solution meets the constraints). Qual represents the trajectory quality score, which can be normalized to [0,1]. . For coefficients, These represent the scores for the safety dimension, smoothness dimension, and efficiency dimension, respectively. . . . This indicates the minimum distance between the updated current driving trajectory and all obstacles. This indicates the preset safe distance threshold. This indicates the maximum acceleration of the current driving trajectory after the update. This indicates the maximum permissible comfortable acceleration. This represents the average speed of the current driving trajectory after the update. Indicates the desired speed.

[0088] When the current iteration is the first iteration, the current sample set is the preset initial sample set. When the current iteration is not the first iteration, the current sample set can be the sample set updated in the previous iteration. The calculated performance metrics can be used as the evaluation values ​​of the target subset; the target subset and the evaluation values ​​can be added to the current sample set as a new subsample set to update the current sample set.

[0089] The current sample set can include one or more subsets. Each subset includes a value for each penalty weight in the penalty weight set, and a corresponding evaluation value. The subset with the largest evaluation value can be selected as the optimal subset from the current sample set. The values ​​in the optimal subset can be the optimal values. The performance index of the target subset can be compared with the evaluation value of the optimal subset; if the performance index is greater than or equal to the evaluation value of the optimal subset, the target subset can be selected as the new optimal subset; if the performance index is less than the evaluation value of the optimal subset, the optimal subset can remain unchanged.

[0090] Therefore, the sample set, the surrogate model, and the optimal subsample set can all be updated during the iteration process. After the iteration ends, the driving trajectory corresponding to the optimal subsample set (the updated current driving trajectory) can be obtained.

[0091] Preset conditions may include iteration termination conditions. For example, preset conditions may include reaching a preset number of iterations.

[0092] If the solver fails to solve the problem or the solution does not meet the constraints, the penalty weights under the safety dimension can be added to the target subset to obtain an adjusted target subset. The adjusted target subset includes adjusted candidate values. The objective function construction steps and optimization steps can be repeated based on the adjusted candidate values ​​of the penalty weights. For example, a linear penalty weight ω under the safety dimension can be added. safe2 The value of ω can be increased by adding ω. safe2 The value of can elevate the safety dimension to a higher priority, forcing the solver to prioritize safety constraints such as obstacle avoidance and lane keeping, which helps improve the solver's convergence stability and the feasibility of the solution. The process of repeatedly executing the objective function construction steps and optimizing the solution steps will not be elaborated further.

[0093] Optionally, if the number of failures reaches a preset number, the configuration mode can be changed from learning mode to degradation mode. This allows the acquisition of preset calibration values. Based on these preset calibration values, the objective function construction and optimization steps can be repeatedly executed to obtain the driving trajectory corresponding to the calibration values.

[0094] Optionally, driving feature data and target subsets can be stored in a corresponding manner to form a cache library. In this way, when similar scenarios are encountered later (e.g., driving feature similarity ≥ 0.8), the penalty weight value can be directly called, skipping the Bayesian optimization iteration and improving real-time performance.

[0095] The cache library can include multiple target subsets and their corresponding driving feature data. Therefore, in the learning mode, the similarity between the current driving feature data and the driving feature data of each target subset in the cache library can be calculated first; the driving feature data with the highest similarity can be selected. If this highest similarity is greater than or equal to a certain threshold, the target subset corresponding to this highest similarity driving feature data can be directly selected, and an objective function can be constructed based on the selected target subset. The objective function can then be optimized using a solver to obtain the solution. If the highest similarity is less than the threshold, the optimal value of the penalty weight can be searched online using a Bayesian optimization algorithm, and the driving trajectory corresponding to the optimal value can be obtained.

[0096] Optionally, offline optimization can be performed on a target subset of driving feature data in the cache periodically (e.g., when the vehicle is charging). Optionally, the driving feature data in the cache can also be subjected to clustering such as K-means clustering to achieve scene segmentation.

[0097] In some embodiments, the configuration mode is either normal mode or extreme mode.

[0098] The penalty weights can be matched to a parameter set based on vehicle driving data. The parameter set may include multiple subsets, and each subset includes the values ​​of the penalty weights in the penalty weight set.

[0099] For example, penalty weights can be matched to a parameter set based on driving feature data. The parameter set can include multiple subsets, each subset corresponding to feature data and including the values ​​of each penalty weight in the penalty weight set. Specifically, for example, the similarity between the driving feature data and the feature data corresponding to each subset can be calculated; the feature data with the highest similarity can be selected, and the subset corresponding to the selected feature data can be obtained.

[0100] The parameter set includes a first parameter set and a second parameter set. The first parameter set includes multiple first subsets, each containing values ​​for quadratic penalty weights and linear penalty weights (e.g., values ​​for quadratic and linear penalty weights within the first, second, third, and fourth piecewise penalty functions). In these first subsets, the values ​​of the quadratic penalty weights are greater than the values ​​of the linear penalty weights. For example, in the first subset, for quadratic and linear penalty weights belonging to the same piecewise penalty function, the value of the quadratic penalty weight is greater than the value of the linear penalty weight. The second parameter set includes multiple second subsets, each containing values ​​for quadratic and linear penalty weights (e.g., values ​​for quadratic and linear penalty weights within the first, second, third, and fourth piecewise penalty functions). In these second subsets, the values ​​of the quadratic penalty weights are less than the values ​​of the linear penalty weights. For example, in the second subset, for quadratic penalty weights and linear penalty weights belonging to the same piecewise penalty function, the value of the quadratic penalty weight is less than the value of the linear penalty weight.

[0101] When the configuration mode is normal, the values ​​of the penalty weights can be matched in the first parameter set. Specifically, the similarity between the driving feature data and the feature data corresponding to each first subset in the first parameter set can be calculated; the feature data with the highest similarity can be selected, and the first subset corresponding to the selected feature data can be obtained. An objective function can be constructed based on the first subset. For example, the current driving trajectory can be obtained; an objective function can be constructed based on the current driving trajectory, the first subset, the switching thresholds of each penalty method, and the values ​​of the weights of each optimization dimension. The objective function can be optimized and solved using a solver to obtain the solution result. The process of constructing the objective function based on the first subset can refer to the process of constructing the objective function based on the candidate values ​​of the penalty weights, which will not be repeated here. The process of optimizing and solving the objective function using a solver to obtain the solution result can refer to the aforementioned embodiments.

[0102] When the configuration mode is in extreme mode, the values ​​of the penalty weights can be matched in the second parameter set. Specifically, the similarity between the driving feature data and the feature data corresponding to each second subset in the second parameter set can be calculated; the feature data with the highest similarity can be selected, and the second subset corresponding to the selected feature data can be obtained. An objective function can be constructed based on the second subset. For example, the current driving trajectory can be obtained; an objective function can be constructed based on the current driving trajectory, the second subset, the switching thresholds of each penalty method, and the values ​​of the weights of each optimization dimension. The objective function can be optimized and solved using a solver to obtain the solution result. The process of constructing the objective function based on the second subset can refer to the process of constructing the objective function based on the candidate values ​​of the penalty weights, which will not be repeated here. The process of optimizing and solving the objective function using a solver to obtain the solution result can refer to the aforementioned embodiments.

[0103] The solution result can include the updated current driving trajectory. The updated current driving trajectory can be a sub-trajectory with one or more time steps added to the current driving trajectory.

[0104] Therefore, in the normal mode, the value of the quadratic penalty weight is greater than the value of the linear penalty weight, so that the quadratic penalty dominates and prioritizes trajectory smoothness. In the extreme mode, the value of the quadratic penalty weight is less than the value of the linear penalty weight, so that the linear penalty dominates and prioritizes solution stability and safety.

[0105] In some embodiments, the configuration mode is a downgrade mode.

[0106] Preset calibration values ​​can be configured for each penalty weight in the penalty weight set; an objective function can be constructed based on the preset calibration values. For example, the current driving trajectory can be obtained; an objective function can be constructed based on the current driving trajectory, preset calibration values, switching thresholds for each penalty method, and the values ​​of the weights for each optimization dimension. The objective function can be optimized and solved using a solver to obtain the solution result. The process of constructing the objective function based on the preset calibration values ​​is similar to the process of constructing the objective function based on the candidate values ​​of the penalty weights, and will not be repeated here. The process of optimizing and solving the objective function using a solver to obtain the solution result is similar to the aforementioned embodiments.

[0107] The solution result can include the updated current driving trajectory. The updated current driving trajectory can be a sub-trajectory with one or more time steps added to the current driving trajectory.

[0108] The vehicle can travel according to the updated driving trajectory.

[0109] Please see Figure 4 This specification provides a vehicle trajectory planning device, comprising the following units.

[0110] The first determining unit 31 is used to determine the configuration mode of the penalty weight based on the vehicle driving data; The second determining unit 32 is used to determine the value of the penalty weight according to the configuration mode; Construction unit 33 is used to construct an objective function based on the value of the penalty weight. The objective function is used to measure the quality of the driving trajectory through multiple segmented penalty functions with multiple optimization dimensions. The multiple function branches of the segmented penalty function correspond to multiple penalty methods. The solver unit 34 is used to optimize the objective function using a solver to obtain the driving trajectory.

[0111] This specification also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described vehicle trajectory planning method.

[0112] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described vehicle trajectory planning method.

[0113] This specification also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described vehicle trajectory planning method.

[0114] Those skilled in the art will understand that this specification can be provided as a method, system, or computer program product. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0115] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments thereof. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. The computer may be a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0116] The functional units in the embodiments of this specification can be integrated into one processing unit, or each functional unit can exist physically separately, or two or more functional units can be integrated into one processing unit.

[0117] Those skilled in the art will understand that the descriptions of the various embodiments in this specification have different focuses, and parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. Furthermore, it is understood that those skilled in the art, after reading this specification, can conceive of any combination of some or all of the embodiments listed in this specification without creative effort, and such combinations are also within the scope of disclosure and protection of this specification.

[0118] Although this specification has been described through embodiments, those skilled in the art will understand that the above embodiments are merely illustrative of the core ideas of this specification. Those skilled in the art will appreciate that many variations and modifications are possible with this specification. It is intended that the appended claims encompass these variations and modifications without departing from the spirit of this specification.

Claims

1. A method for planning vehicle driving trajectory, characterized in that, include: The configuration mode for penalty weighting is determined based on vehicle driving data; Based on the configuration mode, determine the value of the penalty weight; Based on the value of the penalty weight, an objective function is constructed. The objective function is used to measure the quality of the driving trajectory through multiple piecewise penalty functions with multiple optimization dimensions. The multiple function branches of the piecewise penalty function correspond to multiple penalty methods. The objective function is optimized and solved using a solver to obtain the driving trajectory.

2. The method according to claim 1, characterized in that, The configuration mode for determining the penalty weight includes: Identify vehicle driving scenarios based on vehicle driving data; Select the configuration mode that matches the vehicle driving scenario.

3. The method according to claim 2, characterized in that, The identification of vehicle driving scenarios includes: When the vehicle driving data meets the preset rules, the scene corresponding to the preset rules is obtained as the vehicle driving scene; When the vehicle driving data does not meet the preset rules, the vehicle driving data is predicted by a machine learning model to obtain the vehicle driving scenario output by the machine learning model.

4. The method according to claim 3, characterized in that, The selection of a configuration mode that matches the vehicle driving scenario includes: When the confidence level of the vehicle driving scenario is less than or equal to a preset confidence threshold, the configuration mode is determined to be a downgraded mode. The penalty weight is set to a preset calibration value.

5. The method according to claim 1, characterized in that, The method further includes: Calculate the penalty method switching threshold based on the vehicle driving data; The objective function to be constructed includes: Based on the values ​​of the penalty weights and the threshold for switching the penalty method, a target function is constructed.

6. The method according to claim 1, characterized in that, The piecewise penalty function includes a first function branch corresponding to the quadratic penalty method and a second function branch corresponding to the linear penalty method; The first function branch includes a quadratic penalty weight; The second function branch includes linear penalty weights.

7. The method according to claim 6, characterized in that, The piecewise penalty function is: ; Let i represent the piecewise penalty function for the i-th optimization dimension. This is the first function branch. This is the second function branch. This represents the threshold for switching the penalty method in the i-th optimization dimension. This represents the quadratic penalty weight for the i-th optimization dimension. This represents the linear penalty weight for the i-th optimization dimension. This represents the trajectory error of the driving trajectory in the i-th optimization dimension.

8. The method according to claim 1, characterized in that, The configuration mode is learning mode; The process of determining the penalty weight value based on the configuration mode, constructing an objective function based on the penalty weight value, optimizing and solving the objective function to obtain the driving trajectory includes: Iteratively execute the following steps until the preset conditions are met: Using the Bayesian optimization algorithm, candidate values ​​for the penalty weight are selected; Construct the objective function based on the candidate values ​​of the penalty weights; The objective function is optimized and solved using a solver to obtain the solution result; Based on the solution results, calculate the performance index of the candidate values; Update the optimal value of the penalty weight based on the performance metrics of the candidate values; After the iteration is complete, output the driving trajectory corresponding to the optimal value of the penalty weight.

9. The method according to claim 8, characterized in that, The method further includes: Calculate the performance metric for the candidate values ​​based on at least one of the following: The convergence speed of the solver, the feasibility of the driving trajectory, and the quality of the driving trajectory.

10. The method according to claim 8, characterized in that, The multiple optimization dimensions include a safety dimension, the candidate values ​​include the penalty weight values ​​under the safety dimension, and the solution result is a solution failure; the method further includes: By increasing the penalty weight under the security dimension, we obtain the adjusted candidate values; Based on the adjusted candidate values ​​of the penalty weights, the objective function construction steps and optimization solution steps are repeated.

11. The method according to claim 10, characterized in that, The method further includes: If the number of failures reaches the preset number of failures, the configuration mode will be changed from learning mode to degraded mode. Based on the preset values ​​of the penalty weights, the objective function construction steps and optimization solution steps are repeated.

12. The method according to claim 1, characterized in that, The configuration mode is either normal mode or extreme mode; The determination of the penalty weight value includes: Based on vehicle driving data, the penalty weight values ​​are matched in the parameter set. The parameter set includes multiple subsets, and the subsets include the values ​​of the penalty weights.

13. The method according to claim 12, characterized in that, The configuration mode is the normal mode; The piecewise penalty function includes a first function branch corresponding to a quadratic penalty method and a second function branch corresponding to a linear penalty method. The first function branch includes quadratic penalty weights, and the second function branch includes linear penalty weights. The parameter set includes a first parameter set, which includes multiple first subsets. The first subsets include the values ​​of quadratic penalty weights and linear penalty weights. In the first subsets, the value of the quadratic penalty weight is greater than the value of the linear penalty weight.

14. The method according to claim 12, characterized in that, The configuration mode described is an extreme mode; The piecewise penalty function includes a first function branch corresponding to a quadratic penalty method and a second function branch corresponding to a linear penalty method. The first function branch includes quadratic penalty weights, and the second function branch includes linear penalty weights. The parameter set includes a second parameter set, which includes multiple second subsets. The second subsets include the values ​​of quadratic penalty weights and linear penalty weights. In the second subsets, the value of the quadratic penalty weight is less than the value of the linear penalty weight.

15. A vehicle trajectory planning device, characterized in that, include: The first determining unit is used to determine the configuration mode of the penalty weight based on the vehicle driving data; The second determining unit is used to determine the value of the penalty weight according to the configuration mode; The construction unit is used to construct an objective function based on the value of the penalty weight. The objective function is used to measure the quality of the driving trajectory through multiple piecewise penalty functions with multiple optimization dimensions. The multiple function branches of the piecewise penalty function correspond to multiple penalty methods. The solution unit is used to optimize and solve the objective function using a solver to obtain the driving trajectory.

16. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the vehicle trajectory planning method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Automatic driving method and device based on multi-dimensional reward function

    CN120396996A

  • Internet of vehicles safety protection method based on self-learning multi-mode perception

    CN120499665A