Method and apparatus for optimizing decision actions of an autonomous vehicle

CN122546624APending Publication Date: 2026-08-11CHANGCHUN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-09
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]本申请提供一种自动驾驶车辆的决策动作优化方法及装置,以解决相关技术采用各向同性安全评估,导致安全评估有局限,轨迹不平滑,难以满足商业化落地需求的问题

Benefits of technology

[0012]可选地,在本申请的一个实施例中,所述根据所述贝塞尔曲线和所述环境显著性特征向量得到所述当前驾驶场景的动作序列,包括:根据所述环境显著性特征向量生成所述当前驾驶场景的预测动作序列;根据所述预测动作序列生成所述当前驾驶场景的预测轨迹;计算所述预测轨迹和所述贝塞尔曲线之间的均方误差损失;根据所述均方误差损失得到所述动作序列。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122546624A_ABST
    Figure CN122546624A_ABST
Patent Text Reader

Abstract

This application relates to the field of autonomous driving technology, and in particular to a method and apparatus for optimizing decision-making actions of autonomous vehicles. The method includes: acquiring the autonomous vehicle's own state and the neighboring environment states of neighboring vehicles in the current driving scenario; extracting the environmental saliency feature vector of the current driving scenario; constructing the Bézier curve of the current driving scenario; calculating the instantaneous energy risk of the current driving scenario; obtaining the action sequence of the current driving scenario; and obtaining an optimization objective function. The action sequence is then iteratively optimized using the optimization objective function to obtain an optimized action sequence that meets preset optimization conditions. This solves the problem that related technologies employ isotropic safety assessment, leading to limitations in safety assessment, uneven trajectories, and difficulty in meeting the needs of commercialization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving technology, and in particular to a method and apparatus for optimizing the decision-making actions of autonomous vehicles. Background Technology

[0002] As a core application of artificial intelligence, the key to the large-scale commercialization of autonomous driving technology lies in whether the decision-making system can have sufficient safety, reliability and robustness in complex and dynamic traffic environments. The complexity of traffic interactions places higher demands on the performance of autonomous driving decision-making systems, and also drives the continuous iteration and optimization of related technologies.

[0003] In related technologies, autonomous driving systems often adopt a modular design, decoupling perception, prediction, planning, and control into independent sub-tasks. This architecture has the advantage of strong interpretability, but it is highly dependent on manually designed rules. With the opening of large-scale datasets and the improvement of computing power, end-to-end algorithm frameworks have become a research hotspot. Among them, decision-making methods based on deep reinforcement learning can directly output planning trajectory instructions using sensor data, realizing joint optimization of perception and planning features, and performing well in handling complex interaction scenarios.

[0004] However, in related technologies, safety assessment indicators often use isotropic indicators based on Euclidean distance, ignoring the anisotropy in vehicle motion. Furthermore, the continuous motion space of vehicles is large, the convergence speed of end-to-end reinforcement learning algorithms is slow, and the smoothness of trajectory instructions is poor. This makes it difficult to assess high-risk dynamic risks, obtain smooth driving trajectories, reduce driving safety, reduce passenger satisfaction, and fail to meet the needs of large-scale commercialization. Therefore, these issues urgently need to be addressed. Summary of the Invention

[0005] This application provides a method and apparatus for optimizing the decision-making actions of autonomous vehicles, in order to solve the problem that related technologies use isotropic safety assessment, which leads to limited safety assessment, uneven trajectory, and difficulty in meeting the needs of commercialization.

[0006] The first aspect of this application provides a method for optimizing decision-making actions of an autonomous vehicle, comprising the following steps: obtaining the vehicle state of the autonomous vehicle and the neighboring environment states of neighboring vehicles in the current driving scenario; extracting environmental saliency feature vectors of the current driving scenario based on the vehicle state and the neighboring environment states, constructing a Bézier curve of the current driving scenario based on the vehicle state, and calculating the instantaneous energy risk of the current driving scenario based on the vehicle state and the neighboring environment states; obtaining the action sequence of the current driving scenario based on the Bézier curve and the environmental saliency feature vectors, and obtaining an optimization objective function based on the instantaneous energy risk, so as to use the optimization objective function to perform gradient iterative optimization on the action sequence to obtain an optimized action sequence that meets preset optimization conditions.

[0007] Based on the above technical means, this application embodiment obtains the vehicle state and the neighboring environment state, extracts significant environmental features, accurately captures driving scene interaction information, constructs a third-order Bézier curve as a geometric constraint, obtains a predicted action sequence, and calculates instantaneous energy risk to construct an optimization objective function for gradient iterative optimization, thereby obtaining an optimized action sequence. By utilizing the geometric constraints of the Bézier curve and the safety constraints of instantaneous energy risk, sudden action changes are avoided, ensuring that the generated action sequence takes into account both driving safety and comfort, and meets the needs of commercialization of autonomous driving.

[0008] Optionally, in one embodiment of this application, the step of extracting the environmental saliency feature vector of the current driving scene based on the vehicle state and the neighboring environment state includes: concatenating the vehicle state and the neighboring environment state to form a neighboring sequence of the current driving scene; and obtaining the environmental saliency feature vector based on the neighboring sequence.

[0009] Based on the above technical means, the embodiments of this application construct a unified neighbor sequence by splicing the autonomous vehicle state and the neighbor environment state, which can preserve the interaction relationship between the autonomous vehicle and the neighbor vehicles. Then, by generating environmental saliency feature vectors, it can focus on and represent the effective interaction relationship in the current driving scenario, filter redundant information, and improve the compactness and information content of the state representation.

[0010] Optionally, in one embodiment of this application, constructing the Bézier curve of the current driving scenario based on the vehicle state includes: determining the starting point, heading alignment point, target point, and target alignment point of the current driving scenario based on the vehicle state; and constructing the Bézier curve based on the starting point, the heading alignment point, the target point, and the target alignment point.

[0011] Based on the above technical means, the embodiments of this application determine the starting point, heading alignment point, target point and target alignment point of the current driving scenario, and construct a Bézier curve so that the constructed Bézier curve fits the current heading and target heading of the autonomous vehicle, ensuring that the subsequently generated predicted trajectory is smooth and continuous, and improving the smoothness and feasibility of decision-making.

[0012] Optionally, in one embodiment of this application, obtaining the action sequence of the current driving scenario based on the Bézier curve and the environmental saliency feature vector includes: generating a predicted action sequence of the current driving scenario based on the environmental saliency feature vector; generating a predicted trajectory of the current driving scenario based on the predicted action sequence; calculating the mean squared error loss between the predicted trajectory and the Bézier curve; and obtaining the action sequence based on the mean squared error loss.

[0013] Based on the above technical means, the embodiments of this application calculate the mean square error loss between the predicted trajectory and the Bézier curve, and obtain the action sequence based on the mean square error loss. By using the Bézier curve as a supervision signal, the space for exploring invalid actions can be significantly reduced, the network convergence speed can be improved, and the output action sequence can be continuous and comfortable.

[0014] Optionally, in one embodiment of this application, the step of calculating the instantaneous energy risk of the current driving scenario based on the vehicle state and the neighboring environment state includes: calculating the anisotropic safety distance, approach speed, and virtual mass of the current driving scenario based on the vehicle state and the neighboring environment state; and calculating the instantaneous energy risk based on the anisotropic safety distance, the approach speed, and the virtual mass.

[0015] Based on the above technical means, the embodiments of this application can accurately characterize the potential collision risks of traffic participants at different directions and speeds by calculating anisotropic safety distance, approach speed and virtual mass, and calculate instantaneous energy risk based on the above physical quantities, which can objectively reflect the degree of danger of the scene, provide a quantitative risk basis for subsequent online correction, enable the decision system to identify dangerous situations more sensitively and robustly, and improve the active safety performance of autonomous driving.

[0016] Optionally, in one embodiment of this application, the expression of the optimization objective function is: , in, The optimization objective function is... For losses due to rigid safety constraints, For the energy constraint loss of the physical field, For Bessel manifold geometry alignment loss, To control the loss of motion smoothness, To compensate for the loss of motion fidelity, The weighting coefficients for hard safety constraint losses. The weighting coefficients for the energy constraint loss of the physical field. These are the weighting coefficients for the Bessel manifold geometry alignment loss. Weighting coefficients are used to control the loss of motion smoothness.

[0017] Based on the above technical means, the embodiments of this application can integrate hard safety constraint loss, physical field energy constraint loss, Bezier manifold geometric alignment loss, control action smoothness loss, and action fidelity loss by constructing an optimization objective function. By minimizing this objective function during the optimization process, the output optimized action sequence can simultaneously meet the requirements of collision safety, risk avoidance, driving smoothness, and execution feasibility, thereby achieving a synergistic unity of safety, comfort, and decision stability.

[0018] A second aspect of this application provides a decision-making action optimization device for an autonomous vehicle, comprising: an acquisition module for acquiring the vehicle state of the autonomous vehicle and the neighboring environment states of neighboring vehicles in a current driving scenario; a generation module for extracting environmental saliency feature vectors of the current driving scenario based on the vehicle state and the neighboring environment states, constructing a Bézier curve of the current driving scenario based on the vehicle state, and calculating the instantaneous energy risk of the current driving scenario based on the vehicle state and the neighboring environment states; and an optimization module for obtaining an action sequence of the current driving scenario based on the Bézier curve and the environmental saliency feature vectors, obtaining an optimization objective function based on the instantaneous energy risk, and using the optimization objective function to perform gradient iterative optimization on the action sequence to obtain an optimized action sequence that meets preset optimization conditions.

[0019] Based on the above technical means, this application embodiment obtains the vehicle state and the neighboring environment state, extracts significant environmental features, accurately captures driving scene interaction information, constructs a third-order Bézier curve as a geometric constraint, obtains a predicted action sequence, and calculates instantaneous energy risk to construct an optimization objective function for gradient iterative optimization, thereby obtaining an optimized action sequence. By utilizing the geometric constraints of the Bézier curve and the safety constraints of instantaneous energy risk, sudden action changes are avoided, ensuring that the generated action sequence takes into account both driving safety and comfort, and meets the needs of commercialization of autonomous driving.

[0020] A third aspect of this application provides a vehicle, including: a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the program to implement the decision-making action optimization method for an autonomous vehicle as described in the above embodiments.

[0021] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for optimizing decision-making actions of an autonomous vehicle.

[0022] A fifth aspect of this application provides a computer program product, including a computer program that, when executed, implements the above-described method for optimizing decision-making actions of an autonomous vehicle.

[0023] This application's embodiments acquire the vehicle's state and the state of the surrounding environment, extract salient environmental features, accurately capture driving scenario interaction information, and obtain a predicted action sequence by constructing a third-order Bézier curve as a geometric constraint. Instantaneous energy risk is then calculated to construct an optimization objective function for gradient iterative optimization, thereby obtaining an optimized action sequence. By utilizing the geometric constraints of the Bézier curve and the safety constraints of instantaneous energy risk, abrupt action changes are avoided, ensuring that the generated action sequence balances driving safety and comfort, meeting the needs of commercial deployment of autonomous driving. This solves the problem that related technologies use isotropic safety assessments, leading to limited safety assessments, uneven trajectories, and difficulty in meeting commercial deployment requirements.

[0024] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0025] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a decision-making action optimization method for an autonomous vehicle according to an embodiment of this application; Figure 2 This is a block diagram of a decision-making action optimization device for an autonomous vehicle provided according to an embodiment of this application; Figure 3 This is a structural schematic diagram of a vehicle provided according to an embodiment of this application.

[0026] Figure label: 20 - Decision-making action optimization device for autonomous vehicles; 100 - Acquisition module, 200 - Generation module, 300 - Optimization module; 301 - Memory, 302 - Processor, 303 - Communication interface. Detailed Implementation

[0027] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0028] In recent years, with the opening of large-scale datasets and the improvement of computing power, end-to-end algorithm frameworks have gradually become a research hotspot for autonomous driving decision-making. Among them, decision-making methods based on deep reinforcement learning can directly output planned trajectories or control commands using raw sensor data such as images and point clouds, achieving joint optimization of perception and planning features. This data-driven approach, with its powerful feature extraction capabilities and high-dimensional state space processing capabilities, has shown superior application potential compared to related rule-based methods in scenarios such as unstructured roads and complex traffic interactions.

[0029] However, despite the superior performance of deep reinforcement learning methods in closed-loop evaluation, they still face many serious challenges in actual open road deployment. The relevant autonomous driving decision evaluation and planning systems have significant shortcomings in terms of physical interpretability, training efficiency, and real-time safety, specifically in the following three aspects: First, safety assessment metrics suffer from physical deficiencies and limitations. When training and evaluating agents, related technologies often use isotropic metrics based on Euclidean distance (such as simple distance thresholds and collision time) to quantify driving risks. However, such simplified metrics ignore the essential characteristics of vehicle dynamics. On the one hand, vehicle motion is constrained by incompleteness constraints, resulting in a significant anisotropy in risk distribution. For example, longitudinal braking distance is usually much greater than lateral avoidance distance, a physical fact that isotropic metrics cannot accurately reflect. On the other hand, simple distance metrics do not consider the kinetic energy threat posed by relative speed. In high-speed "near miss" scenarios, even if the geometric distance does not reach the safety threshold, the potential collision energy density is already extremely high. Related assessment systems underestimate these high-risk dynamics, leading to a lack of proactive risk avoidance capabilities in highly dynamic traffic interactions.

[0030] Second, high-dimensional space suffers from low efficiency in action exploration and poor trajectory smoothness. End-to-end autonomous driving models need to process multimodal inputs and output continuous control quantities, resulting in a massive action space. In the early stages of reinforcement learning training, the agent lacks effective guidance and can only blindly explore in the vast continuous action space. This not only leads to slow algorithm convergence and low sample efficiency, but also easily generates a large number of kinematically infeasible or extremely oscillating trajectories (such as "snake-like driving" and high-frequency jitter in control commands). Such unsmooth driving behaviors not only increase vehicle mechanical wear, but also severely reduce driving comfort, making it difficult to meet actual driving needs.

[0031] Third, black-box models lack real-time deterministic safety guarantees. Deep neural networks are essentially probabilistic models, which suffer from insufficient interpretability and difficulty in embedding explicit physical rules and constraints. Related end-to-end models typically directly output instantaneous control quantities, lacking explicit constraint mechanisms for future spatiotemporal trajectories. When faced with extreme conditions outside the training distribution, or when the perception module is disturbed (such as domain shift caused by severe weather), pure neural network strategies are prone to outputting dangerous actions that violate traffic rules or even physical laws. Due to the lack of a real-time correction mechanism independent of neural networks and based on physical rules, the system cannot provide a deterministic safety baseline when the model fails, which severely restricts the commercialization of end-to-end autonomous driving technology.

[0032] In summary, there are still many shortcomings in the relevant autonomous driving decision-making technologies. There is an urgent need for an optimization method and device for the decision-making actions of autonomous vehicles that can accurately quantify dynamic driving risks, guide the generation of smooth actions, and provide real-time safety corrections.

[0033] The following description, with reference to the accompanying drawings, describes a method and apparatus for optimizing the decision-making actions of an autonomous vehicle according to embodiments of this application.

[0034] Figure 1 This is a flowchart of a decision-making action optimization method for an autonomous vehicle according to an embodiment of this application.

[0035] like Figure 1 As shown, the method for optimizing the decision-making actions of this autonomous vehicle includes the following steps: In step S101, the autonomous vehicle's status and the neighboring environment status of neighboring vehicles in the current driving scenario are obtained.

[0036] In the embodiments of this application, the current driving scenario can be understood as the set of overall traffic environments in which the autonomous vehicle is located at a certain moment, which may include, but is not limited to, the autonomous vehicle's own state, the neighboring environment state of neighboring vehicles, and the road state, providing scenario conditions for subsequent state acquisition.

[0037] Autonomous vehicles can be understood as the objects that optimize decision-making actions; neighboring vehicles can be understood as traffic participants who have potential interaction risks with the autonomous vehicles.

[0038] In actual implementation, embodiments of this application can acquire status data through sensors (such as LiDAR, cameras). Divided into self-driving state and neighbor's environment Among them, the vehicle status This includes the autonomous vehicle's lateral and longitudinal position, heading angle, speed, acceleration, steering angle, and lateral and heading deviations relative to the global reference path; and the state of the neighboring environment. Including the relative position of each neighboring vehicle ( Relative velocity ( And vehicle type.

[0039] In step S102, the environmental saliency feature vector of the current driving scenario is extracted based on the vehicle state and the neighboring environment state, and the Bézier curve of the current driving scenario is constructed based on the vehicle state. The instantaneous energy risk of the current driving scenario is calculated based on the vehicle state and the neighboring environment state.

[0040] The following details how embodiments of this application extract the environmental saliency feature vector of the current driving scenario based on the vehicle state and the neighboring environment state, construct the Bézier curve of the current driving scenario based on the vehicle state, and calculate the instantaneous energy risk of the current driving scenario based on the vehicle state and the neighboring environment state.

[0041] In one embodiment of this application, the environmental saliency feature vector of the current driving scenario is extracted based on the vehicle state and the neighbor environment state, including: concatenating the vehicle state and the neighbor environment state to form a neighbor sequence of the current driving scenario; and obtaining the environmental saliency feature vector based on the neighbor sequence.

[0042] In actual implementation, the embodiments of this application can be in vehicle mode. and neighbor's environment The sequences are concatenated into a neighbor sequence, which is then input into a Transformer encoder with salient pooling. The Transformer encoder then uses a multi-head attention mechanism to establish the vehicle's state. and neighbor's environment The interaction relationships between them are analyzed to obtain the context embedding vector, and max pooling is performed on the dimension of the context embedding vector to obtain the environmental saliency feature vector of the target dimension. .

[0043] This application embodiment constructs a unified neighbor sequence by splicing the autonomous vehicle state and the neighbor environment state, which can preserve the interaction relationship between the autonomous vehicle and neighbor vehicles. Then, by generating environmental saliency feature vectors, it can focus on and represent the effective interaction relationship in the current driving scenario, filter redundant information, and improve the compactness and information content of the state representation.

[0044] In one embodiment of this application, constructing a Bézier curve for the current driving scenario based on the vehicle's state includes: determining the starting point, heading alignment point, target point, and target alignment point of the current driving scenario based on the vehicle's state; and constructing a Bézier curve based on the starting point, heading alignment point, target point, and target alignment point.

[0045] In actual implementation, the Bézier curve P(τ) constructed in this embodiment requires four control points, namely the starting point. Heading Alignment Point Target alignment point Target point The four control points are defined as follows.

[0046] Starting point The local coordinate origin [0, 0] is the current coordinate origin of the autonomous vehicle, used to ensure that the predicted trajectory is continuously connected with the current position.

[0047] Heading Alignment Point Starting point Extend the preset alignment length along the current heading The point is used to ensure that the tangent direction of the predicted trajectory at the starting point is consistent with the current velocity direction of the autonomous vehicle, satisfying the vehicle non-integrity constraint. Heading alignment point. The expression can be, but is not limited to: , in, For heading alignment point, Starting point For preset alignment length, This represents the current heading angle of the autonomous vehicle.

[0048] Target point These are local waypoints on the navigation path, serving as the endpoints of the predicted trajectory.

[0049] Target alignment point For target point Extend the preset alignment length in the opposite direction of the target heading. The target alignment point is used to ensure that the tangent direction of the predicted trajectory at the endpoint is consistent with the target route, allowing the vehicle to smoothly enter the target lane. The expression can be, but is not limited to: , in, For the target alignment point, For the target point, For preset alignment length, This represents the expected heading angle of the autonomous vehicle at a local waypoint.

[0050] This application embodiment determines the starting point, heading alignment point, target point, and target alignment point of the current driving scenario, and constructs a Bézier curve. The constructed Bézier curve fits the current heading and target heading of the autonomous vehicle, ensuring that the subsequently generated predicted trajectory is smooth and continuous, and improving the smoothness and feasibility of decision-making.

[0051] In one embodiment of this application, calculating the instantaneous energy risk of the current driving scenario based on the vehicle state and the neighboring environment state includes: calculating the anisotropic safe distance, approach speed, and virtual mass of the current driving scenario based on the vehicle state and the neighboring environment state; and calculating the instantaneous energy risk based on the anisotropic safe distance, approach speed, and virtual mass.

[0052] In actual implementation, due to the different hazard sensitivities of autonomous vehicles in the longitudinal and lateral directions, the embodiments of this application first determine the relative positions of each neighboring vehicle ( Converted to anisotropic safe distance Anisotropic safety distance The expression can be, but is not limited to: , in, For anisotropic safety distance, This represents the longitudinal relative distance between the autonomous vehicle and its neighboring vehicles. This refers to the lateral relative distance between the autonomous vehicle and neighboring vehicles. This is the vertical scaling factor. This is the horizontal scaling factor. When... At that time, calculate and The set of values ​​can be used to construct an elliptical risk perception boundary. This application's embodiments define... , making right The contribution of the vehicle is compressed to a greater extent, reflecting the physical characteristic that the longitudinal braking distance of the vehicle is greater than the lateral avoidance distance. Under the same physical distance, the longitudinal hazard sensitivity is higher.

[0053] This application's embodiment calculates the approach speed of neighboring vehicles. This is used to quantify the degree of threat posed by the relative motion of the two vehicles. Approach speed The expression can be, but is not limited to: , in, To approach the speed, Let be the relative position vector between the autonomous vehicle and its neighboring vehicles. Let be the relative velocity vector between the autonomous vehicle and its neighboring vehicles. It is a very small positive number. When this occurs, it indicates that the two vehicles are approaching each other and there is a risk of collision; at this point, the calculation... ; This indicates that the two vehicles are either far apart or stationary, and there is no risk of collision. .

[0054] This application's embodiments calculate the approaching speed of neighboring vehicles, avoiding the problem of only considering distance and ignoring speed. When facing a neighboring vehicle approaching at high speed, even if the neighboring vehicle is at a relatively far distance, its approaching speed is very high, which will trigger a risk warning.

[0055] Furthermore, to demonstrate the destructive force of high-speed objects, embodiments of this application calculate the virtual mass of neighboring vehicles. Virtual quality It's not a fixed value, but the speed of neighboring vehicles. The function. It should be noted that the speed of neighboring vehicles... This refers to the absolute speed of neighboring vehicles in the global coordinate system, and the speed of neighboring vehicles. The higher the value, the greater the destructive power of the neighboring vehicle, reflecting the virtual quality. The larger the approach speed of the neighboring vehicles. It is the relative approach speed of the neighboring vehicle to the autonomous vehicle, and the driving speed of the neighboring vehicle. Different. Virtual quality The expression can be, but is not limited to: , in, The virtual quality of neighboring vehicles. For the physical mass of the neighbor's vehicle, The speed of the neighboring vehicle. This is the first empirical coefficient (used to represent the weight of velocity). The second empirical coefficient ( (used to ensure that risk energy grows superlinearly with speed). This is the third empirical coefficient (used to ensure a reasonable risk value when neighboring vehicles are traveling at low speeds).

[0056] Therefore, the embodiments of this application can calculate the instantaneous risk energy of a single neighboring vehicle to an autonomous vehicle. Instantaneous risk energy The calculation formula can be, but is not limited to, the following: , in, For instantaneous risk energy, For virtual quality, To approach the speed, For anisotropic safety distance, To maximize the scope of influence, For the obstacle scale function, The attenuation index is denoted as . This occurs if and only if the anisotropic safety distance is less than the maximum influence range, i.e. Only then will instantaneous risk energy be calculated. .

[0057] It is understandable that the instantaneous risk energy of a single neighboring vehicle to an autonomous vehicle... It is obtained by multiplying the kinetic term, potential term, and barrier scale term. The kinetic term is... , and virtual quality and approach speed The energy is proportional to the square of the kinetic energy of the collision; the potential energy term is... Safety distance from anisotropy The decay function is inversely proportional to the potential energy; the closer the distance, the higher the potential energy. The barrier scale term is... It increases sharply near the collision boundary, providing a strong gradient and prompting hazard avoidance. (Barrier scale term) The expression can be, but is not limited to: , in, For the obstacle scale term, For the function to find the maximum value, For anisotropic safety distance, The safe threshold distance. When hour, ;when hour, It increased dramatically.

[0058] This application embodiment can accurately characterize the potential collision risks of traffic participants at different orientations and speeds by calculating anisotropic safety distance, approach speed, and virtual mass. Instantaneous energy risk can be calculated based on the above physical quantities, which can objectively reflect the degree of danger of the scene and provide a quantitative risk basis for subsequent online correction. This enables the decision-making system to identify dangerous situations more sensitively and robustly, thereby improving the active safety performance of autonomous driving.

[0059] In step S103, the action sequence of the current driving scenario is obtained based on the Bézier curve and the environmental saliency feature vector, and the optimization objective function is obtained based on the instantaneous energy risk. The action sequence is then iteratively optimized using the optimization objective function to obtain an optimized action sequence that meets the preset optimization conditions.

[0060] In the embodiments of this application, the action sequence can be understood as the predicted action sequence output by the decision network based on the environmental saliency feature vector. By introducing Bézier curves as geometric manifold constraints, the parameters of the decision network are updated, so that the final set of action instructions for the future T steps output by the decision network is used to characterize the continuous driving intention of the autonomous vehicle in the future.

[0061] Preset optimization conditions can be understood as pre-defined termination rules when performing gradient iteration optimization on an action sequence. These rules may include, but are not limited to, the maximum number of gradient iterations and the convergence threshold of the loss function. They are used to limit the computational cost of online correction and ensure the real-time performance of the optimization process.

[0062] Optimizing the action sequence can be understood as taking the action sequence as the initial value and obtaining a set of action instructions for the next T steps under the condition of satisfying the preset optimization. This optimized action sequence retains the original driving intention and meets the requirements of smoothness, comfort and safety, and is used for subsequent controller execution.

[0063] First, this application describes in detail how the embodiments of this application obtain the action sequence of the current driving scenario based on the Bézier curve and the environmental saliency feature vector.

[0064] In one embodiment of this application, obtaining the action sequence of the current driving scenario based on the Bézier curve and the environmental saliency feature vector includes: generating a predicted action sequence of the current driving scenario based on the environmental saliency feature vector; generating a predicted trajectory of the current driving scenario based on the predicted action sequence; calculating the mean squared error loss between the predicted trajectory and the Bézier curve; and obtaining the action sequence based on the mean squared error loss.

[0065] It is understood that the policy network, as a neural network capable of learning the mapping relationship between states and actions, constitutes the bridge from state to action in this application. In the embodiments of this application, the environmental saliency feature vector is input into the policy network, and the corresponding predicted action sequence is output (such as the predicted action sequence for the next T steps, T=5).

[0066] In actual execution, to avoid abrupt changes in actions and uneven trajectories, the embodiments of this application can be based on differentiable programming logic. First, the predicted action sequence is forward extrapolated into a predicted trajectory using a differentiable kinematic model. Then, a Bézier curve is introduced as a geometric manifold constraint. The mean square error loss between the predicted trajectory and the Bézier curve is calculated. The decision network parameters are updated through backpropagation of the mean square error loss, so that the action sequence output by the decision network continuously approaches a smooth and feasible geometric manifold.

[0067] This application embodiment calculates the mean squared error loss between the predicted trajectory and the Bézier curve, and obtains the action sequence based on the mean squared error loss. By using the Bézier curve as a supervision signal, it can significantly reduce the exploration space of invalid actions, improve the network convergence speed, and make the output action sequence continuous and comfortable.

[0068] Then, the embodiments of this application are described in detail how to obtain an optimization objective function based on instantaneous energy risk, and how to use the optimization objective function to perform gradient iterative optimization on the action sequence to obtain an optimized action sequence that meets the preset optimization conditions.

[0069] Based on the description of other embodiments, it can be understood that the embodiments of this application may introduce Bézier curves as geometric manifold constraints, so that the action sequence output by the decision network continuously approaches a smooth and feasible geometric manifold. This process is to punish the disordered exploration actions in the early stage and force the decision network to learn to output smooth actions that conform to the geometric manifold.

[0070] Then, in the inference phase of this application embodiment, when facing normal operating conditions (such as an autonomous vehicle driving at a constant speed and a neighboring vehicle following normally), the decision network outputs a smooth predicted action sequence, and the OCP (Online Correction Planning) module does not make significant modifications to determine the optimized action sequence and send it to the controller. However, when facing extreme and sudden operating conditions (such as a neighboring vehicle suddenly changing lanes at close range into the autonomous vehicle's driving path), considering that the hard safety constraints and physical field energy constraints of the OCP module will generate a large gradient, the predicted trajectory will be forcibly pushed towards the safe zone. If the Bezier manifold geometric alignment constraints are not retained, it is easy to generate an optimized action sequence that the vehicle dynamics cannot execute (such as an instantaneous right-angle turn or a violent zigzag turn) in pursuit of absolute safety.

[0071] Therefore, in the inference phase of this application embodiment, the OCP module calculates the gradient of the objective function with respect to the action sequence through multiple iterations (such as 10 iterations), updates the action sequence according to the gradient, finds the optimal action sequence, and passes the optimal action sequence to the controller for execution to achieve real-time control of the autonomous vehicle.

[0072] In one embodiment of this application, the expression for the optimization objective function is: , in, To optimize the objective function, For losses due to rigid safety constraints, For the energy constraint loss of the physical field, For Bessel manifold geometry alignment loss, To control the loss of motion smoothness, To compensate for the loss of motion fidelity, The weighting coefficients for hard safety constraint losses. The weighting coefficients for the energy constraint loss of the physical field. These are the weighting coefficients for the Bessel manifold geometry alignment loss. Weighting coefficients are used to control the loss of motion smoothness.

[0073] Understandable Used to ensure that actions do not violate safety boundaries. Used to provide repulsive potential energy over long distances, guiding vehicles away from high-risk areas. Used for motion space pruning to ensure that the corrected motion remains on a smooth manifold. This is designed to prevent high-frequency shaking during movement and improve riding comfort. Used to ensure that the modified actions do not deviate from the original strategy intent.

[0074] Specifically, hard safety constraint losses Based on the CBF (Control Barrier Function), a strong penalty gradient is generated by the ReLU activation function when the CBF value of the predicted trajectory point is lower than the safety threshold, forcing the trajectory to back to the safe region.

[0075] Physical field energy constraint loss The instantaneous risk energy calculated above is used to provide a gentle repulsive force before the CBF is triggered, guiding autonomous vehicles to actively move away from high-risk areas at long distances.

[0076] motion fidelity loss This is used to constrain the modified action from deviating from the original policy intent, and its expression can be, but is not limited to, as: , in, To compensate for the loss of motion fidelity, The weights for longitudinal acceleration, This is the corrected longitudinal acceleration. The original longitudinal acceleration, As the weight of the lateral steering angle, This is the corrected lateral steering angle. This represents the original lateral steering angle. Since the loss function represents a penalty cost, a larger weight indicates a higher cost; therefore, this embodiment employs an asymmetric weight design. This indicates that once the optimizer modifies the longitudinal acceleration, the loss value will spike dramatically. In order to minimize the total loss, the optimizer avoids adjusting the longitudinal acceleration and prioritizes adjusting the lateral steering angle, which aligns with the driving intuition that "steering is more comfortable and effective than hard braking."

[0077] Bessel manifold geometric alignment loss Used for action space pruning, ensuring that the corrected action remains on a smooth manifold, its expression can be, but is not limited to: , in, For Bessel manifold geometry alignment loss, To predict the time, For the time step index within the prediction time, This is a differentiable dynamics model used to generate the action sequence output by the decision network. This is converted into the position coordinates of the autonomous vehicle in the global coordinate system. This marks the initial moment of the current plan. The action sequence output by the decision network. For timestamps, For the third-order Bézier curve at the timestamp Spatial coordinates of the points.

[0078] Control motion smoothness loss The sum of squared differences between action sequences at adjacent time steps can be expressed as, but is not limited to, the following: , in, To control the loss of motion smoothness, To predict the time, For the time step index within the prediction time, For the first A sequence of actions at each time step. For the first A sequence of actions at each time step.

[0079] This application embodiment integrates hard safety constraint loss, physical field energy constraint loss, Bezier manifold geometric alignment loss, control action smoothness loss, and action fidelity loss by constructing an optimization objective function. During the optimization process, by minimizing this objective function, the output optimized action sequence can simultaneously meet the requirements of collision safety, risk avoidance, driving smoothness, and execution feasibility, thereby achieving a synergistic unity of safety, comfort, and decision stability.

[0080] Based on the above embodiments, to verify the effectiveness of the autonomous vehicle decision-making action optimization method proposed in this application, a systematic test was conducted on the CARLA high-fidelity simulation platform. The experimental scenarios covered unsignalized intersections, roundabouts, and dense traffic flow, comprehensively simulating various working conditions in actual road driving. The PPO algorithm and the Lagrangian-PPO algorithm were used as comparison objects to quantitatively evaluate the safety, traffic efficiency, and driving comfort of this application. The results show that, compared with the PPO algorithm and the Lagrangian-PPO algorithm, the performance advantages of this application are specifically reflected in the following: (1) Significantly improved safety: The vehicle collision rate was reduced from 32.5% in the comparison algorithm to 1.8%, effectively avoiding various potential collision risks and greatly improving the safety of autonomous driving; (2) Good traffic efficiency: During the experiment, the average vehicle speed was maintained at 7.8 m / s, which was basically the same as the unconstrained baseline speed. There was no phenomenon of vehicle stopping due to excessive pursuit of safety, which ensured normal traffic efficiency. (3) Driving comfort is significantly improved: The lateral acceleration of the vehicle is reduced to 1.25 m / s³, which effectively suppresses the body sway caused by sudden changes in movement, making the vehicle's driving trajectory smoother and more natural, and improving driving comfort.

[0081] The following describes, in the specific scenario of crossing an unsignalized intersection on an urban road, the decision-making action optimization method for autonomous vehicles proposed in this application. In this scenario, the autonomous vehicle (hereinafter referred to as "the vehicle") needs to pass through an unsignalized intersection from east to west, while neighboring vehicles are driving normally in the north-south direction of the intersection. The specific implementation steps are as follows.

[0082] First, this embodiment of the application uses sensing devices such as lidar, camera, and millimeter-wave radar mounted on the vehicle to collect real-time data on the vehicle's status and the surrounding environment. The vehicle's status includes its current global coordinates (x=120 m, y=80 m), heading angle φ=90° (vehicle facing due west), speed v=6.5 m / s, and longitudinal acceleration a=0.3 m / s². The surrounding environment includes the global coordinates, distance relative to the vehicle, speed, and direction of movement of two neighboring vehicles traveling north-south at the intersection. Specifically, neighboring vehicle 1 (coordinates x=120 m, y=95 m, speed v=7 m / s) is traveling south-to-south, and neighboring vehicle 2 (coordinates x=120 m, y=65 m, speed v=6 m / s) is traveling south-to-north. Neither vehicle is changing lanes or undergoing sudden acceleration or deceleration.

[0083] Furthermore, in this embodiment, the acquired vehicle state and neighboring environmental states are concatenated to form a neighbor sequence for the current driving scenario. This neighbor sequence is then processed by a feature extraction network to filter out invalid environmental information and extract a 128-dimensional environmental saliency feature vector. This vector primarily represents the core interaction information such as the relative position and relative speed between the vehicle and two neighboring vehicles. Additionally, the starting point (120 m, 80 m) of the Bézier curve is determined based on the vehicle's current coordinates (x=120 m, y=80 m), and the heading alignment point (125 m, 80 m) is determined based on the vehicle's heading angle. The target point and target alignment point are determined by combining the preset navigation target point (140 m, 80 m) at the intersection exit and the target heading angle (90°). A third-order Bézier curve is constructed based on these four feature points as the geometric manifold constraint for subsequent action sequences. Furthermore, anisotropic safety distances are calculated based on the relative positions of the vehicle and neighboring vehicles (the safety distance between the vehicle and neighboring vehicle 1 is 15 m, and the safety distance with neighboring vehicle 2 is 14 m). m), combined with the relative approach speed of the two vehicles and the virtual mass (the virtual mass of neighboring vehicles is set to 1.0 according to the vehicle size), the instantaneous energy risk value of the current driving scenario is calculated to be 0.32 (risk value range 0-1, the smaller the value, the lower the risk).

[0084] Furthermore, in this embodiment, the extracted environmental saliency feature vector is input into a policy network, which outputs a predicted action sequence for the next 5 steps (T=5, with a time interval of 0.1s per step). This sequence consists of the lateral steering angle and longitudinal acceleration at each moment. The initial predicted action sequence is as follows: Step 1: Steering angle 0°, acceleration 0.2 m / s²; Step 2: Steering angle 0°, acceleration 0.2 m / s²; Step 3: Steering angle 0°, acceleration 0.1 m / s²; Step 4: Steering angle 0°, acceleration 0.1 m / s²; Step 5: Steering angle 0°, acceleration 0.1 m / s²; m / s², and furthermore, based on the calculated instantaneous energy risk, combined with the control motion smoothness loss, motion fidelity loss, and Bézier curve geometric constraint loss, a total optimization objective function is constructed, where the instantaneous energy risk corresponds to the safety penalty term; using this initial predicted motion sequence as the initial value, gradient iterative optimization is performed using the OCP module (the preset optimization condition is a maximum of 10 iterations). In each iteration, the gradient of the optimization objective function with respect to the motion sequence is calculated, and the motion sequence is slightly modified based on the gradient. After 10 iterations, an optimized motion sequence that meets the preset optimization conditions is obtained: Step 1: Turning angle 0°, acceleration 0.2 m / s²; Step 2: Turning angle 0°, acceleration 0.15 m / s²; Step 3: Turning angle 0°, acceleration 0.1 m / s²; Step 4: Turning angle 0°, acceleration 0.05 m / s²; Step 5: Turning angle 0°, acceleration 0.05 m / s²; Step 5: Turning angle 0°, acceleration 0.05 m / s²; Step 6: Turning angle 0°, acceleration 0.05 m / s²; Step 7: Turning angle 0°, acceleration 0.05 m / s²; Step 8: Turning angle 0°, acceleration 0.05 m / s²; Step 9: Turning angle 0°, acceleration 0.05 m / s²; Step 10: Turning angle 0°, acceleration 0.05 m / s²; Step 11: Turning angle 0°, acceleration 0.05 m / s²; Step 12: Turning angle 0°, acceleration 0.05 m / s²; Step 13: Turning angle 0°, acceleration 0.05 m / s²; Step 14: Turning angle 0°, acceleration 0.05 m / s²; Step 15: Turning angle 0°, acceleration 0.05 m / s² The optimized action sequence retains the original decision intention of the vehicle to pass through the intersection in a straight line, while also satisfying safety risk constraints and smoothness constraints. Finally, the OCP module outputs the first action of the optimized action sequence (steering angle 0°, acceleration 0.2 m / s²) to the vehicle controller, and the execution of this action achieves smooth passage.

[0085] The decision-making action optimization method for autonomous vehicles proposed in this application obtains the vehicle's state and the state of the surrounding environment, extracts salient environmental features, accurately captures driving scenario interaction information, constructs a third-order Bézier curve as a geometric constraint, obtains a predicted action sequence, and calculates instantaneous energy risk to construct an optimization objective function for gradient iterative optimization, thereby obtaining an optimized action sequence. By utilizing the geometric constraints of the Bézier curve and the safety constraints of instantaneous energy risk, sudden action changes are avoided, ensuring that the generated action sequence balances driving safety and comfort, meeting the needs of commercialization of autonomous driving. This solves the problem that related technologies use isotropic safety assessment, leading to limited safety assessment, uneven trajectories, and difficulty in meeting commercialization requirements.

[0086] Next, referring to the accompanying drawings, a decision-making action optimization device for an autonomous vehicle according to an embodiment of this application is described.

[0087] Figure 2This is a block diagram of a decision-making action optimization device for an autonomous vehicle provided according to an embodiment of this application.

[0088] like Figure 2 As shown, the decision-making action optimization device 20 for the autonomous vehicle includes: an acquisition module 100, a generation module 200, and an optimization module 300.

[0089] The acquisition module 100 is used to acquire the autonomous vehicle status of the autonomous vehicle and the neighboring environment status of neighboring vehicles in the current driving scenario.

[0090] The generation module 200 is used to extract the environmental saliency feature vector of the current driving scenario based on the vehicle state and the neighboring environment state, construct the Bézier curve of the current driving scenario based on the vehicle state, and calculate the instantaneous energy risk of the current driving scenario based on the vehicle state and the neighboring environment state.

[0091] The optimization module 300 is used to obtain the action sequence of the current driving scenario based on the Bézier curve and the environmental saliency feature vector, and to obtain the optimization objective function based on the instantaneous energy risk. The optimization objective function is used to perform gradient iterative optimization on the action sequence to obtain an optimized action sequence that meets the preset optimization conditions.

[0092] Optionally, in one embodiment of this application, the generation module 200 includes: a splicing unit and a first generation unit.

[0093] The splicing unit is used to splice the vehicle state and the neighbor environment state to form the neighbor sequence of the current driving scenario.

[0094] The first generation unit is used to obtain the environmental saliency feature vector based on the neighbor sequence.

[0095] Optionally, in one embodiment of this application, the generation module 200 includes: a determining unit and a building unit.

[0096] The determining unit is used to determine the starting point, heading alignment point, target point, and target alignment point of the current driving scenario based on the vehicle's status.

[0097] The building unit is used to construct a Bézier curve based on the starting point, heading alignment point, target point, and target alignment point.

[0098] Optionally, in one embodiment of this application, the optimization module 300 includes: a second generation unit, a third generation unit, a first calculation unit, and a fourth generation unit.

[0099] The second generation unit is used to generate a predicted action sequence for the current driving scenario based on the environmental saliency feature vector.

[0100] The third generation unit is used to generate the predicted trajectory of the current driving scene based on the predicted action sequence.

[0101] The first calculation unit is used to calculate the mean square error loss between the predicted trajectory and the Bézier curve.

[0102] The fourth generation unit is used to obtain the action sequence based on the mean squared error loss.

[0103] Optionally, in one embodiment of this application, the generation module 200 includes a second computing unit and a third computing unit.

[0104] The second calculation unit is used to calculate the anisotropic safe distance, approach speed, and virtual mass of the current driving scenario based on the vehicle's status and the status of the neighboring environment.

[0105] The third calculation unit is used to calculate instantaneous energy risk based on anisotropic safety distance, approach velocity, and virtual mass.

[0106] Optionally, in one embodiment of this application, the expression for the objective function can be, but is not limited to, as: , in, To optimize the objective function, For losses due to rigid safety constraints, For the energy constraint loss of the physical field, For Bessel manifold geometry alignment loss, To control the loss of motion smoothness, To compensate for the loss of motion fidelity, The weighting coefficients for hard safety constraint losses. The weighting coefficients for the energy constraint loss of the physical field. These are the weighting coefficients for the Bessel manifold geometry alignment loss. Weighting coefficients are used to control the loss of motion smoothness.

[0107] It should be noted that the foregoing explanation of the method for optimizing the decision-making actions of autonomous vehicles also applies to the device for optimizing the decision-making actions of autonomous vehicles in this embodiment, and will not be repeated here.

[0108] The decision-making action optimization device for autonomous vehicles proposed in this application obtains the vehicle's state and the state of the surrounding environment, extracts salient environmental features, accurately captures driving scene interaction information, and obtains a predicted action sequence by constructing a third-order Bézier curve as a geometric constraint. It also calculates instantaneous energy risk to construct an optimization objective function for gradient iterative optimization, thereby obtaining an optimized action sequence. By utilizing the geometric constraints of the Bézier curve and the safety constraints of instantaneous energy risk, it avoids abrupt action changes, ensuring that the generated action sequence balances driving safety and comfort, meeting the needs of commercialization of autonomous driving. This solves the problem that related technologies use isotropic safety assessment, leading to limited safety assessments, uneven trajectories, and difficulty in meeting commercialization requirements.

[0109] Figure 3 This is a schematic diagram of the structure of a vehicle according to an embodiment of this application. The vehicle may include: The memory 301, the processor 302, and the computer program stored on the memory 301 and capable of running on the processor 302.

[0110] When the processor 302 executes the program, it implements the decision-making action optimization method for autonomous vehicles provided in the above embodiments.

[0111] Furthermore, the vehicle also includes: Communication interface 303 is used for communication between memory 301 and processor 302.

[0112] The memory 301 is used to store computer programs that can run on the processor 302.

[0113] The memory 301 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0114] If the memory 301, processor 302, and communication interface 303 are implemented independently, then the communication interface 303, memory 301, and processor 302 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0115] Optionally, in a specific implementation, if the memory 301, processor 302, and communication interface 303 are integrated on a single chip, then the memory 301, processor 302, and communication interface 303 can communicate with each other through an internal interface.

[0116] Processor 302 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0117] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described method for optimizing the decision-making actions of an autonomous vehicle.

[0118] This application also provides a computer program product, including a computer program that, when executed, implements the above-described method for optimizing decision-making actions of an autonomous vehicle.

[0119] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0120] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0121] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0122] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0123] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or more of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0124] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0125] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0126] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for optimizing decision-making actions of an autonomous vehicle, characterized in that, Includes the following steps: Obtain the autonomous vehicle status and the neighboring environment status of neighboring vehicles in the current driving scenario; The environmental saliency feature vector of the current driving scenario is extracted based on the vehicle state and the neighboring environment state, and the Bezier curve of the current driving scenario is constructed based on the vehicle state. The instantaneous energy risk of the current driving scenario is calculated based on the vehicle state and the neighboring environment state. The action sequence of the current driving scenario is obtained based on the Bézier curve and the environmental saliency feature vector, and an optimization objective function is obtained based on the instantaneous energy risk. The action sequence is then subjected to gradient iterative optimization using the optimization objective function to obtain an optimized action sequence that meets the preset optimization conditions.

2. The method of claim 1, wherein, The step of extracting the environmental saliency feature vector of the current driving scenario based on the vehicle state and the neighboring environment state includes: The vehicle state and the neighbor environment state are spliced ​​together to form the neighbor sequence of the current driving scenario; The environmental saliency feature vector is obtained based on the neighbor sequence.

3. The method of claim 1, wherein, The step of constructing the Bezier curve for the current driving scenario based on the vehicle's state includes: The starting point, heading alignment point, target point, and target alignment point of the current driving scenario are determined based on the vehicle's status. The Bézier curve is constructed based on the starting point, the heading alignment point, the target point, and the target alignment point.

4. The method of claim 1, wherein, The step of obtaining the action sequence of the current driving scenario based on the Bézier curve and the environmental saliency feature vector includes: Generate a predicted action sequence for the current driving scenario based on the environmental saliency feature vector; Generate a predicted trajectory for the current driving scenario based on the predicted action sequence; Calculate the mean squared error loss between the predicted trajectory and the Bézier curve; The action sequence is obtained based on the mean squared error loss.

5. The method of claim 1, wherein, The calculation of the instantaneous energy risk of the current driving scenario based on the vehicle's state and the neighboring environment state includes: Calculate the anisotropic safe distance, approach speed, and virtual mass of the current driving scenario based on the vehicle's state and the neighbor's environmental state; The instantaneous energy risk is calculated based on the anisotropic safety distance, the approach speed, and the virtual mass.

6. The method of claim 1, wherein, The expression for the optimization objective function is: , in, The optimization objective function is... For losses due to rigid safety constraints, For the energy constraint loss of the physical field, For Bessel manifold geometry alignment loss, To control the loss of motion smoothness, To compensate for the loss of motion fidelity, The weighting coefficients for hard safety constraint losses. The weighting coefficients for the energy constraint loss of the physical field. These are the weighting coefficients for the Bessel manifold geometry alignment loss. Weighting coefficients are used to control the loss of motion smoothness.

7. A decision action optimization device of an autonomous vehicle, characterized by, include: The acquisition module is used to acquire the autonomous vehicle status of the autonomous vehicle and the neighboring environment status of neighboring vehicles in the current driving scenario. The generation module is used to extract the environmental saliency feature vector of the current driving scenario based on the vehicle state and the neighboring environment state, construct the Bézier curve of the current driving scenario based on the vehicle state, and calculate the instantaneous energy risk of the current driving scenario based on the vehicle state and the neighboring environment state. The optimization module is used to obtain the action sequence of the current driving scenario based on the Bézier curve and the environmental saliency feature vector, and to obtain an optimization objective function based on the instantaneous energy risk, so as to use the optimization objective function to perform gradient iterative optimization on the action sequence to obtain an optimized action sequence that meets the preset optimization conditions.

8. A vehicle characterized by comprising: include: A memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the program to implement the method of optimizing a decision action of an autonomous vehicle as claimed in any of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor for implementing the method of optimizing a decision action of an autonomous vehicle as claimed in any of claims 1-6.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed for implementing the method of optimizing a decision action of an autonomous vehicle as claimed in any of claims 1-6.