Intelligent vehicle risk-sensitive sequential behavior decision method, device and equipment
By constructing a dynamic objective function and rolling time-domain optimization, the problem of multi-objective coordination and multi-stage stable decision-making for intelligent vehicles in complex scenarios is solved, enabling intelligent vehicles to drive safely and efficiently in real-world environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2023-06-29
- Publication Date
- 2026-05-05
AI Technical Summary
In existing technologies, intelligent vehicle decision-making methods have requirements in terms of training data sample size and quality, making them difficult to apply to real-world complex dynamic scenarios. This leads to instability in multi-objective collaboration and multi-stage decision-making, and a lack of effective risk-sensitive sequential behavior decision-making methods.
By acquiring the driving status information of traffic participants in a preset traffic environment, a dynamic objective function is constructed. Based on the longitudinal and lateral dynamic safety margins, the vehicle's single-walking behavior decision-making strategy is determined, the driving intentions of surrounding vehicles are identified, the cost of the behavioral decision-making strategy is calculated, and the strategy is adjusted through rolling time-domain optimization until the risk-sensitive sequential decision-making strategy is consistent with the vehicle's actions, and the optimal trajectory is output.
It enables dynamic multi-objective collaboration and multi-stage stable decision-making of intelligent vehicles in complex scenarios, promoting the application and development of intelligent vehicles.
Smart Images

Figure CN116572993B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent vehicle application technology, and in particular to a method, apparatus and equipment for risk-sensitive sequential behavior decision-making for intelligent vehicles. Background Technology
[0002] In complex scenarios, intelligent vehicle decision-making systems need to output stable, continuous, and reasonable decision strategies to meet actual driving needs. However, during actual vehicle operation, performance requirements are diverse and coupled, and the multi-stage, multi-scenario decision-making process is not coherent, posing numerous challenges to research on multi-objective collaboration and multi-stage decision performance assurance for intelligent vehicles. The key challenge lies in how to complete sequential behavioral decisions for vehicles in dynamic environments, planning feasible trajectories to meet performance goals such as safety, efficiency, and reliability. In the human-vehicle-road system, the driver plays multiple roles, including decision-maker and experiencer. The driver's perception-decision-control characteristics directly affect the vehicle's handling stability and safety. Drivers have a general cognitive mechanism and common control patterns regarding potential risks in the traffic environment, but different types of risk sources have varying impacts on drivers. The driver's risk response affects their scenario understanding and decision strategy output, thereby impacting driving safety.
[0003] In related technologies, research on advanced driver assistance systems (ADAS) and autonomous driving aims to improve the intelligence level of vehicles by selecting the most appropriate driving behavior to adapt to complex and dynamic traffic environments. However, in the process of driving a vehicle, the ability of a driver to effectively cope with the complex and ever-changing traffic environment under the coupled effects of multiple factors such as people, vehicles, and roads is not based on a single dangerous scenario, but rather on the coordinated unity of "perception-assessment-decision" in any scenario. Therefore, it is necessary to draw inspiration from the proactive decision-making methods adopted by drivers in dealing with complex and ever-changing traffic environments, so that autonomous driving systems can adapt to open and dynamic traffic scenarios and achieve safe, reliable, agile, and smooth autonomous driving in real traffic.
[0004] Furthermore, given the randomness of traffic participant behavior in complex dynamic interactive environments (including the randomness of intention and trajectory interactions) and the dynamic nature of traffic environment states (static / dynamic blind spots, sensor errors, etc.), research on decision-making systems for intelligent vehicles is crucial. During vehicle operation, these systems need to output optimal and stable decision-making strategies to meet actual driving needs. Currently, there is a wealth of research on intelligent vehicle decision-making both domestically and internationally.
[0005] In related technologies, centralized decision-making frameworks are mainly based on an integrated approach. They learn or explore driving behavior data through end-to-end methods (such as deep learning and reinforcement learning) based on environmental information received by sensors. Based on the input information from the sensor end, they directly output vehicle-level control commands.
[0006] In related technologies, hierarchical decision-making frameworks decompose the entire decision-making process into a series of sub-functional modules, each of which can be designed independently. Typically, behavioral decisions are made first, followed by trajectory planning. The entire decision-making process under a hierarchical framework can be categorized into single-stage or single-step behavioral decision-making, and multi-stage sequential or multi-step behavioral decision-making. Single-step behavioral decision-making methods include traditional rule-based / optimization-based methods, probabilistic statistical reasoning-based methods, and behavioral interaction-based methods. Multi-stage or multi-step sequential decision-making (or trajectory planning) methods mainly include search, interpolation, sampling, and artificial potential fields.
[0007] Among the aforementioned technologies, data-driven centralized decision-making methods do not rely on limited expert rules for decision-making. Their policy networks directly generate test cases from real driving data, resulting in better overall and real-time decision-making. However, centralized decision-making frameworks (such as reinforcement learning-based and supervised learning-based methods) essentially learn from and imitate natural driving data, requiring a certain level of training data sample size and quality, or a certain quantity and quality of test scenarios. In real-world complex dynamic scenarios, it is even more necessary to consider the impact and constraints of traffic rules or road terrain on the vehicle's continuous sequential decision-making. This can greatly simplify the analysis of complex interactive problems and support the vehicle in making reasonable multi-stage decisions. By comparison, traditional hierarchical decision-making frameworks have simple logic and can output better and more stable behavioral decisions, but knowledge acquisition is difficult, and transferring them to different scenarios is challenging. End-to-end solutions can handle multiple driving scenarios, but their interpretability is poor.
[0008] In summary, there is currently a lack of a method for risk-sensitive sequential behavior decision-making in intelligent vehicles, which urgently needs to be addressed. Summary of the Invention
[0009] This application provides a risk-sensitive sequential behavior decision-making method, apparatus, and device for intelligent vehicles to address the problems in related technologies, such as the requirement for training data sample size and quality, making it difficult to apply to real-world complex dynamic scenarios. This method enables intelligent vehicles to achieve dynamic multi-objective collaboration and multi-stage stable decision-making in complex scenarios, thereby promoting the application and development of intelligent vehicles.
[0010] The first aspect of this application provides a risk-sensitive sequential behavior decision-making method for intelligent vehicles, including the following steps:
[0011] Obtain driving status information of traffic participants under a preset traffic environment;
[0012] Based on the driving status information, a dynamic objective function is constructed according to the driver's risk sensitivity.
[0013] Based on the dynamic objective function, longitudinal dynamic safety margin, and lateral dynamic safety margin, the vehicle is determined to walk alone as the decision-making strategy.
[0014] Based on the single-walking behavior decision-making strategy, the driving intentions of surrounding vehicles are identified, and the cost of different behavioral decision-making strategies for the current vehicle is calculated based on the driving intentions. The optimal strategy for the current vehicle is then matched based on the cost values.
[0015] The action of the current vehicle at the current moment is determined based on the optimal strategy of the current vehicle. Based on the rolling time domain optimization, it is determined whether the effect of the single-step output behavior decision strategy is consistent with the actual effect of the action output of the current vehicle at the current moment. If they are inconsistent, the driving status information of traffic participants in the preset traffic environment is reacquired until the effect of the risk-sensitive sequential decision strategy output is consistent with the actual effect of the action output of the current vehicle at the current moment. The optimal trajectory is output according to the risk-sensitive sequential decision strategy.
[0016] Optionally, in some embodiments, constructing a dynamic objective function based on the driving state information and the driver's risk sensitivity includes:
[0017] Construct the minimum action quantity in the real physical system, and determine the unified driving objective in the driver's active decision-making process based on the driving state information;
[0018] Based on the unified driving objective, output a dynamic objective function based on the minimum action amount.
[0019] The dynamic objective function is:
[0020]
[0021] Where i represents intelligent vehicle, S Risk Let t0 be the dynamic objective function of intelligent vehicle i during the decision-making and planning process, and t be the initial time. f It is the end time, L i For the Lagrange equation of a two-vehicle system, T i For the vehicle's kinetic energy, U i This represents the system's potential energy.
[0022] Optionally, in some embodiments, before determining the vehicle's one-way walking as the decision strategy based on the dynamic objective function, the longitudinal dynamic safety margin, and the lateral dynamic safety margin, the method further includes:
[0023] Based on the interaction between the vehicle and the traffic participants, the preset traffic environment is divided into multiple dual-vehicle systems composed of two vehicles according to the interaction between the vehicles.
[0024] The Lagrange equations of the dual-vehicle system are determined, and the longitudinal dynamic safety margin and the lateral dynamic safety margin are determined in the decision-making process based on the Lagrange equations of the dual-vehicle system.
[0025] Optionally, in some embodiments, the Lagrange equation for the dual-vehicle system is:
[0026]
[0027] The lateral dynamic safety margin is:
[0028] r y =r ij,y +∈;
[0029] The longitudinal dynamic safety margin is:
[0030] r x =r ij,x +Ψ(v ix ,(v ix -v jx ))Δt;
[0031] Where i represents the vehicle, T i For the vehicle's kinetic energy, U i Let m be the system potential energy. i It is the mass of vehicle i, v i Let v be the speed of vehicle i. j Let t be the speed of vehicle j, t0 be the initial time, and t be the speed of vehicle j. f It is the end time, R i It is the longitudinal constraint resistance of traffic rules on the driver, G i G is the virtual driving force generated by the driver driving the intelligent vehicle. i,x G is the longitudinal target driving force for the driver. i,y For the driver's lateral target driving force, v ix Let v be the longitudinal velocity of vehicle i. iy Let F be the lateral velocity of vehicle i. li, and F li, F represents the lateral constraint force generated by the two lane lines of the lane in which the target vehicle is traveling. ji r represents the interaction risk force exerted by vehicle j on vehicle i; ij, It is the longitudinal following distance between vehicles i and j, r ij, Let ∈ be the following distance between vehicles i and j in the lateral direction, ∈ be the lateral safety margin, Ψ(*) be the positive correlation function, and Δt be the reversal time step.
[0032] Optionally, in some embodiments, the risk-sensitive sequential decision-making strategy includes:
[0033] The vehicle driving strategy is optimized and adjusted continuously within the time window to obtain the optimal dynamic objective function expressed in the rolling time domain.
[0034] The extreme values of the functional are solved using a preset variational method, and the risk-sensitive sequential decision-making strategy is obtained based on the solution results.
[0035] Optionally, in some embodiments, the optimal dynamic objective function is expressed in the rolling time domain as:
[0036]
[0037]
[0038] Where S is the actual action quantity, k represents time, J(·) is the cost function, u(k) is the input vector, x(k) is the state vector, and Φ is the target set; S * For the ideal action, H w Represents the rolling time domain, τ is the time increment, u(k+τ|k) is the control input value from time k to the future time (k+τ), X(k+τ|k) is the predicted value from time k to the future time (k+τ), and Z(·) is the end penalty term.
[0039] A second aspect of this application provides a risk-sensitive sequential behavior decision-making device for intelligent vehicles, comprising:
[0040] The traffic participant driving information acquisition module is used to acquire the driving status information of traffic participants under a preset traffic environment;
[0041] The dynamic objective function construction module is used to construct a dynamic objective function based on the driving state information and the driver's risk sensitivity.
[0042] The intelligent vehicle single-walk behavior decision module is used to determine the vehicle single-walk behavior decision strategy based on the dynamic objective function, longitudinal dynamic safety margin, and lateral dynamic safety margin.
[0043] The behavioral decision cost calculation and strategy selection module is used to identify the driving intentions of surrounding vehicles based on the single-walk behavioral decision strategy, calculate the cost of different behavioral decision strategies adopted by the current vehicle according to the driving intentions, and match the optimal strategy for the current vehicle according to the cost value; and
[0044] The risk-sensitive sequential decision-making model construction and optimization module is used to determine the current vehicle's action at the current moment based on the optimal strategy of the current vehicle, and to determine whether the effect of the single-step output behavior decision strategy is consistent with the actual effect of the current vehicle's action output at the current moment based on rolling time domain optimization. If they are inconsistent, the module reacquires the driving status information of traffic participants in the preset traffic environment until the effect of the risk-sensitive sequential decision-making strategy output is consistent with the actual effect of the current vehicle's action output at the current moment, and outputs the optimal trajectory according to the risk-sensitive sequential decision-making strategy.
[0045] Optionally, in some embodiments, the dynamic objective function construction module is specifically used for:
[0046] Construct the minimum action quantity in the real physical system, and determine the unified driving objective in the driver's active decision-making process based on the driving state information;
[0047] Based on the unified driving objective, output a dynamic objective function based on the minimum action amount.
[0048] The dynamic objective function is:
[0049]
[0050] Where i represents intelligent vehicle, S Risk Let t0 be the dynamic objective function of intelligent vehicle i during the decision-making and planning process, and t be the initial time. f It is the end time, L i For the Lagrange equation of a two-vehicle system, T i For the vehicle's kinetic energy, U i This represents the system's potential energy.
[0051] Optionally, in some embodiments, before determining the vehicle's one-way walking decision strategy based on the dynamic objective function, the longitudinal dynamic safety margin, and the lateral dynamic safety margin, the intelligent vehicle one-way walking decision module is further configured to:
[0052] Based on the interaction between the vehicle and the traffic participants, the preset traffic environment is divided into multiple dual-vehicle systems composed of two vehicles according to the interaction between the vehicles.
[0053] The Lagrange equations of the dual-vehicle system are determined, and the longitudinal dynamic safety margin and the lateral dynamic safety margin are determined in the decision-making process based on the Lagrange equations of the dual-vehicle system.
[0054] Optionally, in some embodiments, the Lagrange equation for the dual-vehicle system is:
[0055]
[0056] The lateral dynamic safety margin is:
[0057] r y =r ij,y +∈;
[0058] The longitudinal dynamic safety margin is:
[0059] r x =r ij,x +Ψ(v ix ,(v ix -v jx ))Δt;
[0060] Where i represents the vehicle, T i For the vehicle's kinetic energy, U i Let m be the system potential energy. i It is the mass of vehicle i, v u Let v be the speed of vehicle i. j Let t be the speed of vehicle j, t0 be the initial time, and t be the speed of vehicle j. f It is the end time, R i It is the longitudinal constraint resistance of traffic rules on the driver, G i G is the virtual driving force generated by the driver driving the intelligent vehicle. i,x G is the longitudinal target driving force for the driver. i,y For the driver's lateral target driving force, v ix Let v be the longitudinal velocity of vehicle i. iy Let F be the lateral velocity of vehicle i. li, and F li, F represents the lateral constraint force generated by the two lane lines of the lane in which the target vehicle is traveling. ji r represents the interaction risk force exerted by vehicle j on vehicle i; ij, It is the longitudinal following distance between vehicles i and j, r ij, Let ∈ be the following distance between vehicles i and j in the lateral direction, ∈ be the lateral safety margin, Ψ(*) be the positive correlation function, and Δt be the reversal time step.
[0061] Optionally, in some embodiments, the risk-sensitivity sequential decision-making strategy includes:
[0062] The vehicle driving strategy is optimized and adjusted continuously within the time window to obtain the optimal dynamic objective function expressed in the rolling time domain.
[0063] The extreme values of the functional are solved using a preset variational method, and the risk-sensitive sequential decision-making strategy is obtained based on the solution results.
[0064] Optionally, in some embodiments, the optimal dynamic objective function is expressed in the rolling time domain as:
[0065]
[0066]
[0067] Where S is the actual action quantity, k represents time, J(·) is the cost function, u(k) is the input vector, x(k) is the state vector, and Φ is the target set; S * For the ideal action, H w Represents the rolling time domain, τ is the time increment, u(k+τ|k) is the control input value from time k to the future time (k+τ), X(k+τ|k) is the predicted value from time k to the future time (k+τ), and Z(·) is the end penalty term.
[0068] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the intelligent vehicle risk-sensitive sequential behavior decision-making method as described in the above embodiments.
[0069] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the intelligent vehicle risk-sensitive sequential behavior decision-making method as described in the above embodiments.
[0070] Therefore, this application obtains the driving state information of traffic participants under a preset traffic environment, constructs a dynamic objective function, determines the vehicle's single-walk decision-making strategy based on the dynamic objective function, determines the longitudinal and lateral dynamic safety margins during the decision-making process, identifies the driving intentions of surrounding vehicles based on the single-walk decision-making strategy, calculates the cost of different behavioral decision-making strategies for the vehicle, matches the vehicle's optimal strategy, and repeats the above steps until the risk-sensitive sequential decision-making strategy is consistent with the vehicle's action at the current moment, and outputs the optimal trajectory according to the risk-sensitive sequential decision-making strategy. This solves the problems in related technologies, such as the requirement for certain training data sample size and quality in intelligent vehicle decision-making methods, making them difficult to apply to real-world complex dynamic scenarios. It enables dynamic multi-objective collaboration and multi-stage stable decision-making for intelligent vehicles in complex scenarios, promoting the application and development of intelligent vehicles.
[0071] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0072] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0073] Figure 1 This is a flowchart of a risk-sensitive sequential behavior decision-making method for intelligent vehicles provided according to an embodiment of this application;
[0074] Figure 2 This is a schematic diagram illustrating the calculation and selection of a single-walk behavior decision-making strategy for an intelligent vehicle according to an embodiment of this application;
[0075] Figure 3 This is a schematic diagram illustrating the rolling execution of multi-walk behavior decisions in intelligent vehicle interaction according to an embodiment of this application;
[0076] Figure 4 This is a schematic diagram of a rolling time-domain optimization strategy for intelligent vehicles according to an embodiment of this application;
[0077] Figure 5 This is a schematic diagram of a risk-sensitive sequential behavior decision-making method for intelligent vehicles according to an embodiment of this application;
[0078] Figure 6 This is a schematic diagram of a risk-sensitive sequential behavior decision-making method for intelligent vehicles according to an embodiment of this application;
[0079] Figure 7 This is a schematic diagram comparing the ideas of single-walk behavior decision-making, driver, and classic decision-making methods in any scenario according to an embodiment of this application;
[0080] Figure 8 This is a schematic diagram of a dynamic multi-objective behavior decision-making process for an intelligent vehicle according to an embodiment of this application;
[0081] Figure 9 This is a schematic diagram of a risk-sensitive sequential behavior decision-making device for intelligent vehicles provided according to an embodiment of this application;
[0082] Figure 10 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0083] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0084] The following description, with reference to the accompanying drawings, outlines a method, apparatus, and device for making risk-sensitive sequential behavior decisions for intelligent vehicles, representing embodiments of this application. To address the issue that intelligent vehicle decision-making methods mentioned in the background technology have certain requirements in terms of training data sample size and quality, making them difficult to apply to real-world complex dynamic scenarios, this application provides a risk-sensitive sequential behavior decision-making method for intelligent vehicles. In this method, the driving state information of traffic participants in a preset traffic environment is acquired; based on the driving state information, a dynamic objective function is constructed according to the driver's risk sensitivity; based on the dynamic objective function, longitudinal dynamic safety margin, and lateral dynamic safety margin, a single-step behavior decision-making strategy for the vehicle is determined; based on the single-step behavior decision-making strategy, the driving intentions of surrounding vehicles are identified, and the cost of different behavior decision-making strategies adopted by the current vehicle is calculated according to the driving intentions, and the optimal strategy for the current vehicle is matched according to the cost value; the action of the current vehicle at the current moment is determined according to the optimal strategy of the current vehicle, and the effect of the single-step output behavior decision-making strategy is judged based on rolling time-domain optimization to see if it is consistent with the actual effect of the current vehicle's action output at the current moment. If inconsistent, the driving state information of traffic participants in the preset traffic environment is reacquired until the effect of the risk-sensitive sequential decision-making strategy output is consistent with the actual effect of the current vehicle's action output at the current moment, and the optimal trajectory is output according to the risk-sensitive sequential decision-making strategy. This solves the problem that intelligent vehicle decision-making methods have certain requirements in terms of training data sample size and quality, making them difficult to apply to real-world complex dynamic scenarios. It enables intelligent vehicles to achieve dynamic multi-objective collaboration and multi-stage stable decision-making in complex scenarios, thus promoting the application and development of intelligent vehicles.
[0085] Specifically, Figure 1 This is a flowchart illustrating a risk-sensitive sequential behavior decision-making method for intelligent vehicles provided in an embodiment of this application.
[0086] like Figure 1 As shown, the intelligent vehicle risk-sensitive sequential behavior decision-making method includes the following steps:
[0087] In step S101, the driving status information of traffic participants under a preset traffic environment is obtained.
[0088] The preset traffic environment refers to a complex traffic environment, or even any scenario. Since real traffic environments include various vehicles, cyclists, and pedestrians, traffic participants in this embodiment refer to the vehicle itself, following vehicles, and surrounding obstacles. Because the intelligent vehicle risk-sensitive sequential behavior decision-making method of this application considers the complexity of driving scenarios, the unpredictability of traffic participant behavior, and the driver's dynamic needs for driving safety, this embodiment requires first obtaining the driving information of traffic participants in a complex traffic environment to make risk-sensitive sequential behavior decisions based on this information.
[0089] In step S102, a dynamic objective function is constructed based on the driving status information and the driver's risk sensitivity.
[0090] It should be noted that after obtaining the driving status information of traffic participants, this application then considers the driver's risk sensitivity to construct a dynamic objective function that unifies safety and efficiency.
[0091] Specifically, this application constructs a mathematical expression for the minimum action in a real physical system, thereby determining a unified driving objective of pursuing safety and efficiency in the driver's active decision-making process, and further determining an output objective function based on the minimum action that considers the driver's risk sensitivity.
[0092] Optionally, in some embodiments, a dynamic objective function is constructed based on driving state information and the driver's risk sensitivity, including: constructing a minimum action in the real physical system, and determining a unified driving objective in the driver's active decision-making process based on driving state information; and outputting a dynamic objective function based on the minimum action based on the unified driving objective, wherein the dynamic objective function is:
[0093]
[0094] Where i represents intelligent vehicle, S Risk Let t0 be the dynamic objective function of intelligent vehicle i during the decision-making and planning process, and t be the initial time. f It is the end time, L i For the Lagrange equation of a two-vehicle system, T i For the vehicle's kinetic energy, U i This represents the system's potential energy.
[0095] Specifically, in a system composed of free point masses, this application, based on the principle of least action, considers the actual motion of the point masses from spatial point 1 to spatial point 2, such that the integral... A minimum value exists. To describe the interaction between particles, this application defines the function generated by their interaction as -. According to the additivity of the Lagrange equation:
[0096]
[0097] Where, r a Let be the radius vector of the a-th particle.
[0098] It is understandable that in real physical systems, theoretically, a path with minimum action from the starting point to the ending point can be found, and this path satisfies Newton's laws. Therefore, this application can establish that the objective function for intelligent vehicles pursuing safety and efficiency can be transformed into finding a path with minimum action S. minThe path. The action S is defined as the integral of the Lagrange L between the two endpoints:
[0099]
[0100]
[0101] Where T represents the vehicle's kinetic energy, U represents the system's potential energy, t0 is the initial time, and t f It is the end time.
[0102] It should be noted that, in determining the unified driving goals of safety and efficiency pursued by the driver during the proactive decision-making process, this application explores and identifies the principles and rules of dynamic interaction between the driver and intelligent vehicle in complex dynamic traffic scenarios. To achieve anthropomorphism in the intelligent vehicle driving process, based on the driver's behavioral decision-making process, it extracts the main goals pursued during the interaction with the environment, focusing on the fundamental goals pursued by the driver during driving: safety and efficiency. Safety is related to multiple factors, including vehicle attributes (such as speed and mass) and interaction characteristics with other vehicles (including relative distance and speed), and is a basic guarantee for the driving process. Efficiency mainly refers to the pursuit of driving speed.
[0103] Furthermore, considering the driving goals of safety and efficiency pursued by drivers, this application establishes a Lagrangian quantity reflecting vehicle safety, while efficiency can be expressed as the time consumed throughout the driving process. Based on the principle of least action, the cost of generating a feasible path can be defined as S. Risk The dynamic objective function of intelligent vehicle i in the decision-making and planning process can be characterized as:
[0104]
[0105] Where i represents intelligent vehicle, S Risk Let t0 be the dynamic objective function of intelligent vehicle i during the decision-making and planning process, and t be the initial time. f It is the end time, L i For the Lagrange equation of a two-vehicle system, T i For the vehicle's kinetic energy, U i This represents the system's potential energy.
[0106] Therefore, this application draws on the principle of "least action" in physics to simulate the cognitive decision-making mechanism of drivers. With the goal of optimizing driving safety and traffic efficiency, it proposes a dynamic multi-performance objective collaborative intelligent vehicle decision-making method, which makes up for the shortcomings of difficulty in weighting multiple single objectives and difficulty in determining the dimensions, and realizes the optimization of intelligent vehicle decision-making in complex road traffic environments.
[0107] In step S103, the vehicle's single-walking decision strategy is determined based on the dynamic objective function, longitudinal dynamic safety margin, and lateral dynamic safety margin.
[0108] It should be noted that, after step S102, this application obtains the dynamic objective function. By calculating, this application can obtain the longitudinal dynamic safety margin and the lateral dynamic safety margin. Therefore, based on the dynamic objective function, the longitudinal dynamic safety margin, and the lateral dynamic safety margin, this application can determine the vehicle's single-walk behavior decision strategy.
[0109] Specifically, Figure 2 This is a schematic diagram illustrating the calculation and selection of a single-walking decision-making strategy for an intelligent vehicle according to an embodiment of this application. Figure 2 As shown, when selecting a single-walk behavior decision strategy based on a unified objective function, the intelligent vehicle interacts with traffic participants during its operation. Given the uncertainty of the surrounding vehicles' behavior and the dynamic and static uncertainties of the environment, the differences in driving goals and dynamic needs among dynamic and static traffic participants (such as surrounding vehicles) necessitate maneuvers such as overtaking, lane changing, and lane keeping during the vehicle's operation. Therefore, during the interaction between the two vehicles, the autonomous vehicle first determines the driving intentions of surrounding vehicles, considering potential risks and driver expectations through the decision objective function. Based on this, this application calculates the cost of different behavior decision strategies adopted by the autonomous vehicle according to the principle of least action, and theoretically, can output a superior and stable behavior decision.
[0110] Optionally, in some embodiments, before determining the vehicle's single-walking decision strategy based on the dynamic objective function, longitudinal dynamic safety margin, and lateral dynamic safety margin, the method further includes: dividing the preset traffic environment into multiple two-vehicle systems composed of two vehicles based on the interaction between the vehicle and traffic participants; determining the Lagrange equation of the two-vehicle system; and determining the longitudinal dynamic safety margin and lateral dynamic safety margin in the decision-making process based on the Lagrange equation of the two-vehicle system.
[0111] Optionally, in some embodiments, the Lagrange equation for the two-vehicle system is:
[0112]
[0113] The lateral dynamic safety margin is:
[0114] r y =r ij,y +∈; (6)
[0115] The longitudinal dynamic safety margin is:
[0116] r x =r ij,x +Ψ(v ix ,(v ix -vjx ))Δt; (7)
[0117] Where i represents the vehicle, T i For the vehicle's kinetic energy, U i Let m be the system potential energy. i It is the mass of vehicle i, v i Let v be the speed of vehicle i. j Let t be the speed of vehicle j, t0 be the initial time, and t be the speed of vehicle j. f It is the end time, R i It is the longitudinal constraint resistance of traffic rules on the driver, G i G is the virtual driving force generated by the driver driving the intelligent vehicle. i,x G is the longitudinal target driving force for the driver. i,y For the driver's lateral target driving force, v ix Let v be the longitudinal velocity of vehicle i. iy Let F be the lateral velocity of vehicle i. li, and F li, F represents the lateral constraint force generated by the two lane lines of the lane in which the target vehicle is traveling. ji r represents the interaction risk force exerted by vehicle j on vehicle i; ij, It is the longitudinal following distance between vehicles i and j, r ij, Let ∈ be the following distance between vehicles i and j in the lateral direction, ∈ be the lateral safety margin, Ψ(*) be the positive correlation function, and Δt be the reversal time step.
[0118] Specifically, based on the interaction between vehicles and traffic participants, this application divides the complex traffic environment into several dual-vehicle systems composed of two vehicles, which can also be divided into simple following systems, cutting-in systems, etc., based on the interaction between vehicles.
[0119] The Lagrange equation for the dual-vehicle system in this application is:
[0120]
[0121] Where i represents the vehicle; m i It is the mass of vehicle i; v i and v j Let G be the speeds of vehicles i and j, respectively. i This is the virtual driving force generated by the driver's target driving the intelligent vehicle, which enables the vehicle to move from the starting position to the ending position. When there is no lane change, the driver's target driving force only exists in the longitudinal direction (G). i,x When choosing to change lanes, due to lateral movement, the driver's target driving force will have a lateral component G. i,yDuring vehicle movement, the longitudinal constraints imposed by road speed limits and the lateral constraints imposed by road markings must be specifically considered. Among these, R... i It is the longitudinal constraint resistance of traffic rules on the driver; F li, and F li, These represent the lateral constraint forces generated by the two lane lines of the lane in which the target vehicle is traveling; F ji This represents the risk force of the interaction between vehicle j and vehicle i.
[0122] Furthermore, during vehicle motion, the vehicle's own speed and its relative speed with surrounding vehicles directly affect the potential collision risk. Generally, the higher the vehicle's speed, the greater the collision risk. The greater the relative speed between the vehicle and surrounding vehicles, the greater the traffic interference and potential impact on vehicles in front and behind. Therefore, the vehicle safety margin in the longitudinal direction is positively correlated with both vehicle speed and relative speed. Moreover, the driver's elliptical field of view and visual attention distribution during driving are also limiting factors. The lateral safety margin is related to the driver's risk sensitivity, and considering the limited variation in lateral speed, the margin can be defined as a variable that is only related to dynamic relative distance. Therefore, this application defines the longitudinal and lateral dynamic safety margins as follows:
[0123] r y =r ij,y +∈; (6)
[0124] r x =r ij,x +Ψ(v ix ,(v ix -v jx ))Δt; (7)
[0125] Where Ψ(*) is the positive correlation function; Δt represents the backoff time step; and ∈ is the lateral safety margin. ij, and r ij, The distribution represents the following distances of vehicles i and j in the longitudinal and lateral directions.
[0126] In step S104, based on the single-walk behavior decision-making strategy, the driving intentions of surrounding vehicles are identified, and the cost of different behavior decision-making strategies adopted by the current vehicle is calculated according to the driving intentions. The optimal strategy of the current vehicle is then matched according to the cost value.
[0127] It should be noted that this application first determines the expression of a single-walk behavioral decision-making model that considers efficiency (kinetic energy manifestation) and safety (potential energy guarantee), then calculates the cost of different behavioral decision-making strategies adopted by the vehicle, and selects an appropriate strategy based on the calculated cost output.
[0128] Specifically, for any driving scenario, assuming there is vehicle i in the traffic system, its Lagrange equation Li It can be described as:
[0129]
[0130] In the process of one-way behavioral decision-making that considers the driver's risk sensitivity, the cost of one-way behavioral decision-making can be characterized as S. Risk :
[0131]
[0132] Where t0 is the initial time, t f It is the end time.
[0133] Furthermore, during the operation of the intelligent vehicle, traffic participants interact with it. Due to differences in driving goals and dynamic needs, both static and dynamic traffic participants (such as surrounding vehicles) require the autonomous vehicle to perform maneuvers such as overtaking, lane changing, and lane keeping. Therefore, during the interaction between the two vehicles, the autonomous vehicle in this embodiment first determines the driving intentions of the surrounding vehicles, then calculates the cost of different behavioral decision-making strategies, and selects the appropriate strategy based on the calculated cost.
[0134] In step S105, the action of the current vehicle at the current moment is determined according to the optimal strategy of the current vehicle, and the effect of the single-step output behavior decision strategy is determined based on the rolling time domain optimization to see if it is consistent with the actual effect of the action output of the current vehicle at the current moment. If they are inconsistent, the driving status information of traffic participants in the preset traffic environment is reacquired until the effect of the risk-sensitive sequential decision strategy output is consistent with the actual effect of the action output of the current vehicle at the current moment. The optimal trajectory is output according to the risk-sensitive sequential decision strategy.
[0135] Those skilled in the art will understand that the solution output during a single-step decision-making process may not be the optimal solution, or even non-existent, leading to decision oscillations in complex scenarios. Therefore, based on the realization of single-step behavioral decision-making for intelligent vehicles, this application, considering that the multi-step sequential behavioral decision-making process is a continuous multi-stage process, and utilizing the structural characteristics of the traffic environment, further optimizes and selects the optimal path after generating feasible finite candidate trajectory curves.
[0136] In some embodiments, this application outputs the vehicle's action at time t based on a single-step decision and executes it. At time t+1, a risk-sensitive sequential decision-making method is constructed based on rolling time-domain optimization. The above steps are repeated until multi-step sequential behavioral decision-making is achieved through rolling execution, and a trajectory to reach the destination safely and efficiently is found.
[0137] Optionally, in some embodiments, the risk-sensitive sequential decision-making strategy includes: continuously optimizing and adjusting the vehicle driving strategy within a time window to obtain the optimal dynamic objective function expressed in the rolling time domain; solving the functional extrema based on a preset variational method, and obtaining the risk-sensitive sequential decision-making strategy based on the solution results.
[0138] Specifically, Figure 3 This is a schematic diagram illustrating the multi-walk action decision-making and rolling execution of intelligent vehicle interaction according to an embodiment of this application, as shown below. Figure 3 As shown, multi-step reasoning decision-making can expand the solution space, and better decisions can be obtained through iterative optimization and rolling execution, ensuring stability and continuity in the time domain, and achieving local optimal or even global optimal decisions. Under the condition of optimal cost, the vehicle will complete the lane-changing maneuver, continue to maintain a safe distance from the vehicle in front, and then recalculate the cost of the next stage of behavior decision-making strategy and select the next stage of lane-changing or following behavior.
[0139] Furthermore, based on the realization of intelligent vehicle single-step behavior decision-making, considering that the multi-step sequential behavior decision-making process is a continuous multi-stage process, and utilizing the structural characteristics of the traffic environment, it is necessary to further optimize and select the optimal path after generating feasible finite candidate trajectory curves. Figure 4 This is a schematic diagram of the intelligent vehicle rolling time-domain optimization strategy according to an embodiment of this application, as shown below. Figure 4 As shown, the rolling time-domain optimization method can achieve optimal control of specific constrained systems, effectively solving for optimal control in real time based on the system's output state and constraints at sampling moments. It employs the idea of dividing the decision-making process into multiple stages to solve multivariable and constrained optimization problems online.
[0140] Additionally, it should be noted that the entire optimization process in this application will output the current best state based on the existing state. The solution is executed in a rolling fashion to ultimately obtain the optimal trajectory, which consists of a series of locally (time-optimal) trajectory segments.
[0141] Optionally, in some embodiments, the optimal dynamic objective function is expressed in the rolling time domain as:
[0142]
[0143]
[0144] Where S is the actual action quantity, k represents time, J(·) is the cost function, u(k) is the input vector, x(k) is the state vector, and Φ is the target set; S * For the ideal action, H wRepresents the rolling time domain, τ is the time increment, u(k+τ|k) is the control input value from time k to the future time (k+τ), X(k+τ|k) is the predicted value from time k to the future time (k+τ), and Z(·) is the end penalty term.
[0145] In some embodiments, the sampling time is set to 50ms, and the rolling time domain H w Setting it to 15, when the rolling time domain is set to a larger value (such as 20), although the output trajectory performance is better, it will result in a longer computation time; if the rolling time domain is set to a smaller value (such as 10), the dynamic prediction range will be insufficient, and the effect may be poor.
[0146] Specifically, this application considers the main idea of rolling time-domain optimization as follows: based on the corresponding objective function and constraints, iterative solutions are performed to finally obtain the optimal input for a finite time period at that moment.
[0147] In this system, the intelligent vehicle, under the conditions of obstacle avoidance and satisfying corresponding constraints, will reach the target area at the minimum cost. Its dynamics, defined as a linear time-invariant system, are described as follows:
[0148] X(k+1)=AX(k)+Bu(k); (12)
[0149]
[0150] Where X(k) is the state vector and u(k) is the input vector. The output state vector satisfies the following constraints: That is, equation (12) is the state space model of the vehicle, and equation (13) is the constraint condition that the state and input need to satisfy.
[0151] In the actual cost calculation process, this application uses the variational method to calculate and solve the S of this path. R1sk The minimum value can obtain the optimal trajectory of a multi-stage process, that is, to solve the functional extremum.
[0152] Specifically, the amount of action S i This is a functional that quantifies the driver's pursuit of multiple driving objectives; the functional S i The extreme value of S represents the extreme value pursued by the driver for multiple objectives. Therefore, this application calculates and solves for the S of this path. Risk The minimum value can obtain the optimal trajectory of a multi-stage process. The variational method is often used in numerical solutions to find the extrema of a functional, that is:
[0153]
[0154] in, It is the theoretically minimum action quantity, and the actual action S of each trajectory. Risk The extreme value.
[0155] Applying the Euler–Lagrange fundamental equations to solve the optimization problem, the expression of the fundamental equations is as follows:
[0156]
[0157] Based on the above embodiments, combined with Figure 5 , Figure 5 This is a schematic diagram of the intelligent vehicle risk-sensitive sequential behavior decision-making method according to an embodiment of this application. The implementation of the entire sequential behavior decision-making process of this application includes:
[0158] 1. Generate and calculate the actual effect of each trajectory;
[0159] 2. Screen for risk-sensitive, safe, collision-free feasible trajectories;
[0160] 3. Determine whether the actual action of the trajectory is equal to or close to the theoretical minimum value;
[0161] 4. Continuous calculation and multiple iterations output the optimal trajectory. Ultimately, this achieves the output of the best decision for the current state over a continuous time series.
[0162] Furthermore, in the sequential behavior decision-making problem defined in this application based on rolling time-domain optimization, the objective function considers a safe, efficient, and optimal equilibrium. The main constraints on the intelligent vehicle's driving process are soft constraints of traffic rules (including lane line constraints, speed limits imposed by laws and regulations, etc.), hard constraints of surrounding traffic participants, and hard constraints of vehicle dynamics, which are specifically reflected in the potential field function and constraint conditions and are considered during the solution process. Finally, through rolling calculation and multiple iterations, the optimal trajectory is output.
[0163] Therefore, the intelligent vehicle risk-sensitive sequential behavior decision-making method proposed in this application, based on a unified objective function, constructs a single-step behavior decision-making and multi-step sequential decision-making method based on rolling time-domain optimization, which ensures stability and continuity in the time domain and realizes the output of the best decision based on the current vehicle state in a continuous time series.
[0164] To enable those skilled in the art to further understand the intelligent vehicle risk-sensitive sequential behavior decision-making method of this application, the following embodiments are provided in conjunction with the accompanying drawings to illustrate the steps of the method.
[0165] Specifically, Figure 6 This is a schematic diagram of the intelligent vehicle risk-sensitive sequential behavior decision-making method according to an embodiment of this application, as shown below. Figure 6 As shown, the method includes the following steps:
[0166] Step S601: Perceive and acquire driving status information of surrounding traffic participants in complex traffic environments;
[0167] Step S602: Construct a dynamic objective function that unifies safety and efficiency, taking into account the driver's risk sensitivity;
[0168] Step S603: Based on the objective function driven in step S602, construct a single-walk behavior decision-making method for intelligent vehicles and determine the longitudinal and lateral dynamic safety margins during the decision-making process.
[0169] Step S604: Based on the single-walk behavior decision-making method, determine the driving intentions of surrounding vehicles, calculate the cost of different behavior decision-making strategies for the vehicle, and select an appropriate strategy based on the calculated cost output.
[0170] Step S605: Output the vehicle's action at time t based on the single-step decision and execute it;
[0171] Step S606: At time t+1, construct a risk-sensitive sequential decision-making method based on rolling time-domain optimization;
[0172] Step S607: Repeat the above steps until multi-step sequential behavioral decision-making is achieved through rolling execution, and a safe and efficient trajectory to the destination is found.
[0173] Therefore, the intelligent vehicle risk-sensitive sequential behavior decision-making method of this application can take into account the complexity of road scenarios, the individual randomness of traffic participants, and the interactive game between individuals, describe the interaction mechanism between traffic elements and the driver's expected driving goals, meet dynamic multi-objective requirements, and thus realize the dynamic multi-objective collaboration and multi-stage stable decision-making of intelligent vehicles in complex scenarios.
[0174] In some embodiments, Figure 7 This diagram illustrates a comparison of the behavioral decision-making approach, driver's approach, and classic decision-making method in any scenario according to embodiments of this application. To verify the rationality of the decision-making logic of the proposed method, the solution approaches of the behavioral decision-making method, driver's approach, and classic decision-making method are compared in any scenario. In the vehicle decision-making process, classic methods are often rule-based, attempting to predetermine how to handle each obstacle. The difficulty lies in the fixed threshold parameters in different scenarios, making it difficult to transfer to arbitrary scenarios and increasing the probability of application failure. For drivers, they acquire driving strategies based on experience, but due to the limitations of their own driving experience and style, differences can also lead to traffic accidents. The behavioral decision-making method proposed in this application is driven by the principle of least action. This method divides the planning process into two stages: feasible behavioral decision generation and decision optimization evaluation. It learns the handling mechanisms of excellent drivers and establishes a comprehensive trajectory quality evaluation function that balances safety and efficiency, achieving an objective expression of the intelligent vehicle trajectory quality evaluation function in different scenarios. This can be used to solve trajectory planning problems in complex environments. Therefore, for intelligent vehicle driving safety evaluation, it can ensure that no collisions occur during optimization.
[0175] Of course, in other embodiments, Figure 8 This diagram illustrates the dynamic multi-objective behavioral decision-making process of an intelligent vehicle according to an embodiment of this application. To address the complexity of driving scenarios, the unpredictability of traffic participant behavior, and the dynamic demands of drivers for driving safety, efficiency, and comfort, this application requires intelligent vehicles to balance multiple performance objectives to achieve optimal performance during trajectory planning in different scenarios. Excellent drivers can effectively cope with complex and uncertain environments, and the intelligent vehicle decision-making system can be compared to the driver's brain. Therefore, how to think like a human, simulate the unified decision-making logic of a driver, and improve the vehicle's intelligence level by learning from the driver, achieving maximum "humanization," is a key challenge in the design of decision-making algorithms. The behavioral decision-making method of this application mines behavioral characteristics from a large amount of natural driving data and combines them with the properties of physical systems in nature. This avoids the bounded rationality and cognitive biases of drivers, thus balancing the dynamic pursuit of different objectives by intelligent vehicles in different scenarios.
[0176] Therefore, this application constructs a risk-sensitive sequential behavior decision-making method for intelligent vehicles, which enables stable, reliable, and continuous decision-making in multiple scenarios and stages for intelligent vehicles, thus promoting the application and development of intelligent vehicles.
[0177] The intelligent vehicle risk-sensitive sequential behavior decision-making method proposed in this application obtains the driving state information of traffic participants in a preset traffic environment, constructs a dynamic objective function, determines the vehicle's single-walking behavior decision strategy based on the dynamic objective function, determines the longitudinal and lateral dynamic safety margins during the decision-making process, identifies the driving intentions of surrounding vehicles based on the single-walking behavior decision strategy, calculates the cost of different behavior decision strategies adopted by the vehicle to match the optimal strategy of the vehicle, and repeats the above steps until the risk-sensitive sequential decision strategy is consistent with the vehicle's action at the current moment, and outputs the optimal trajectory according to the risk-sensitive sequential decision strategy. This solves the problems in related technologies, such as the requirement for certain training data sample size and quality, making it difficult to apply to real-world complex dynamic scenarios. It enables intelligent vehicles to achieve dynamic multi-objective collaboration and multi-stage stable decision-making in complex scenarios, promoting the application and development of intelligent vehicles.
[0178] Next, referring to the accompanying drawings, a risk-sensitive sequential behavior decision-making device for intelligent vehicles according to an embodiment of this application is described.
[0179] Figure 9 This is a block diagram of an intelligent vehicle risk-sensitive sequential behavior decision-making device according to an embodiment of this application.
[0180] like Figure 9As shown, the intelligent vehicle risk-sensitive sequential behavior decision-making device 10 includes: a traffic participant driving information acquisition module 100, a dynamic objective function construction module 200, an intelligent vehicle single-walk behavior decision-making module 300, a behavior decision cost calculation and strategy selection module 400, and a risk-sensitive sequential decision-making model construction and optimization module 500.
[0181] The system includes: a traffic participant driving information acquisition module 100, used to acquire driving status information of traffic participants under a preset traffic environment; a dynamic objective function construction module 200, used to construct a dynamic objective function based on the driving status information and the driver's risk sensitivity; an intelligent vehicle single-walk behavior decision-making module 300, used to determine the vehicle's single-walk behavior decision-making strategy based on the dynamic objective function, longitudinal dynamic safety margin, and lateral dynamic safety margin; and a behavior decision cost calculation and strategy selection module 400, used to identify the driving intentions of surrounding vehicles based on the single-walk behavior decision-making strategy, and calculate the cost of different behavior decision-making strategies adopted by the current vehicle based on the driving intentions. Furthermore, the system matches the optimal strategy for the current vehicle based on the cost value; and a risk-sensitive sequential decision model construction and optimization module 500 is used to determine the current vehicle's action at the current moment based on the current vehicle's optimal strategy, and to determine whether the effect of the single-step output behavior decision strategy is consistent with the actual effect of the current vehicle's action output at the current moment based on the rolling time domain optimization. If they are inconsistent, the system reacquires the driving status information of traffic participants under the preset traffic environment until the effect of the risk-sensitive sequential decision strategy output is consistent with the actual effect of the current vehicle's action output at the current moment, and outputs the optimal trajectory based on the risk-sensitive sequential decision strategy.
[0182] Optionally, in some embodiments, the dynamic objective function construction module 200 is specifically used to: construct the minimum action in the real physical system, and determine the unified driving objective in the driver's active decision-making process based on driving state information; and output a dynamic objective function based on the minimum action based on the unified driving objective, wherein the dynamic objective function is:
[0183]
[0184] Where i represents intelligent vehicle, S Risk Let t0 be the dynamic objective function of intelligent vehicle i during the decision-making and planning process, and t be the initial time. f It is the end time, L i For the Lagrange equation of a two-vehicle system, T i For the vehicle's kinetic energy, U i This represents the system's potential energy.
[0185] Optionally, in some embodiments, before determining the vehicle's single-walking behavior decision-making strategy based on the dynamic objective function, longitudinal dynamic safety margin, and lateral dynamic safety margin, the intelligent vehicle single-walking behavior decision-making module 300 is further configured to: divide the preset traffic environment into multiple dual-vehicle systems composed of two vehicles based on the interaction between the vehicle and traffic participants; determine the Lagrange equation of the dual-vehicle system; and determine the longitudinal dynamic safety margin and lateral dynamic safety margin in the decision-making process based on the Lagrange equation of the dual-vehicle system.
[0186] Optionally, in some embodiments, the Lagrange equation for the two-vehicle system is:
[0187]
[0188] The lateral dynamic safety margin is:
[0189] r y =r ij,y +∈;
[0190] The longitudinal dynamic safety margin is:
[0191] r x =r ij,x +Ψ(v ix ,(v ix -v jx ))Δt;
[0192] Where i represents the vehicle, T i For the vehicle's kinetic energy, U i Let m be the system potential energy. i It is the mass of vehicle i, v i Let v be the speed of vehicle i. j Let t be the speed of vehicle j, t0 be the initial time, and t be the speed of vehicle j. f It is the end time, R i It is the longitudinal constraint resistance of traffic rules on the driver, G i G is the virtual driving force generated by the driver driving the intelligent vehicle. i,x G is the longitudinal target driving force for the driver. i,y For the driver's lateral target driving force, v ix Let v be the longitudinal velocity of vehicle i. iy Let F be the lateral velocity of vehicle i. li,1 and F li,2 F represents the lateral constraint force generated by the two lane lines of the lane in which the target vehicle is traveling. ji r represents the interaction risk force exerted by vehicle j on vehicle i; ij, It is the longitudinal following distance between vehicles i and j, r ij,Let ∈ be the following distance between vehicles i and j in the lateral direction, ∈ be the lateral safety margin, Ψ(*) be the positive correlation function, and Δt be the reversal time step.
[0193] Optionally, in some embodiments, the risk-sensitive sequential decision-making strategy includes: continuously optimizing and adjusting the vehicle driving strategy within a time window to obtain the optimal dynamic objective function expressed in the rolling time domain; solving the functional extremum based on a preset variational method, and obtaining the risk-sensitive sequential decision-making strategy based on the solution results.
[0194] Optionally, in some embodiments, the optimal dynamic objective function is expressed in the rolling time domain as:
[0195]
[0196]
[0197] Where S is the actual action quantity, k represents time, J(·) is the cost function, u(k) is the input vector, x(k) is the state vector, and Φ is the target set; S * For the ideal action, H w Represents the rolling time domain, τ is the time increment, u(k+τ|k) is the control input value from time k to the future time (k+τ), X(k+τ|k) is the predicted value from time k to the future time (k+τ), and Z(·) is the end penalty term.
[0198] It should be noted that the foregoing explanation of the embodiment of the intelligent vehicle risk-sensitive sequential behavior decision-making method also applies to the intelligent vehicle risk-sensitive sequential behavior decision-making device of this embodiment, and will not be repeated here.
[0199] The intelligent vehicle risk-sensitive sequential behavior decision-making device proposed in this application acquires the driving state information of traffic participants in a preset traffic environment, constructs a dynamic objective function, determines the vehicle's single-walking behavior decision-making strategy based on the dynamic objective function, determines the longitudinal and lateral dynamic safety margins during the decision-making process, identifies the driving intentions of surrounding vehicles based on the single-walking behavior decision-making strategy, calculates the cost of different behavior decision-making strategies adopted by the vehicle, and matches the vehicle's optimal strategy. The above steps are repeated until the risk-sensitive sequential decision-making strategy matches the vehicle's current action, and the optimal trajectory is output based on the risk-sensitive sequential decision-making strategy. This solves the problems in related technologies, such as the requirement for certain training data sample size and quality in intelligent vehicle decision-making methods, making them difficult to apply to real-world complex dynamic scenarios. It enables dynamic multi-objective collaboration and multi-stage stable decision-making for intelligent vehicles in complex scenarios, promoting the application and development of intelligent vehicles.
[0200] Figure 10A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:
[0201] The memory 1001, the processor 1002, and the computer program stored on the memory 1001 and capable of running on the processor 1002.
[0202] When the processor 1002 executes the program, it implements the risk-sensitive sequential behavior decision-making method for intelligent electronic devices provided in the above embodiments.
[0203] Furthermore, electronic devices also include:
[0204] Communication interface 1003 is used for communication between memory 1001 and processor 1002.
[0205] The memory 1001 is used to store computer programs that can run on the processor 1002.
[0206] The memory 1001 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk storage.
[0207] If the memory 1001, processor 1002, and communication interface 1003 are implemented independently, then the communication interface 1003, memory 1001, and processor 1002 can be interconnected via a bus to complete communication between them. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 10 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0208] Optionally, in a specific implementation, if the memory 1001, processor 1002, and communication interface 1003 are integrated on a single chip, then the memory 1001, processor 1002, and communication interface 1003 can communicate with each other through an internal interface.
[0209] The processor 1002 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of this application.
[0210] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described intelligent vehicle risk-sensitive sequential behavior decision-making method.
[0211] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0212] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0213] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0214] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (FPGAs), field-programmable gate arrays (FPGAs), etc.
[0215] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0216] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A risk-sensitive sequential behavior decision-making method for intelligent vehicles, characterized in that, Includes the following steps: Obtain driving status information of traffic participants under a preset traffic environment; Based on the driving status information, a dynamic objective function is constructed according to the driver's risk sensitivity. Based on the dynamic objective function, longitudinal dynamic safety margin, and lateral dynamic safety margin, the vehicle is determined to walk alone as the decision-making strategy. Based on the single-walking behavior decision-making strategy, the driving intentions of surrounding vehicles are identified, and the cost of different behavior decision-making strategies adopted by the current vehicle is calculated according to the driving intentions. The optimal strategy of the current vehicle is then matched according to the cost value. as well as The action of the current vehicle at the current moment is determined based on the optimal strategy of the current vehicle. Based on the rolling time domain optimization, it is determined whether the effect of the single-step output behavior decision strategy is consistent with the actual effect of the action output of the current vehicle at the current moment. If they are inconsistent, the driving status information of traffic participants in the preset traffic environment is reacquired until the effect of the risk-sensitive sequential decision strategy output is consistent with the actual effect of the action output of the current vehicle at the current moment. The optimal trajectory is output according to the risk-sensitive sequential decision strategy.
2. The method according to claim 1, characterized in that, The step of constructing a dynamic objective function based on the driving status information and the driver's risk sensitivity includes: Construct the minimum action quantity in the real physical system, and determine the unified driving objective in the driver's active decision-making process based on the driving state information; Based on the unified driving objective, output a dynamic objective function based on the minimum action amount. The dynamic objective function is: ; in, i For intelligent vehicles, For intelligent vehicles i The dynamic objective function in the decision-making and planning process. It is the initial time. It is the end time. For the Lagrange equations of a two-vehicle system, T i For the vehicle's kinetic energy, U i This represents the system's potential energy.
3. The method according to claim 1, characterized in that, Before determining the vehicle's one-way walking as the decision strategy based on the dynamic objective function, the longitudinal dynamic safety margin, and the lateral dynamic safety margin, the process further includes: Based on the interaction between the vehicle and the traffic participants, the preset traffic environment is divided into multiple dual-vehicle systems composed of two vehicles according to the interaction between the vehicles. The Lagrange equations of the dual-vehicle system are determined, and the longitudinal dynamic safety margin and the lateral dynamic safety margin are determined in the decision-making process based on the Lagrange equations of the dual-vehicle system.
4. The method according to claim 3, characterized in that, The Lagrange equation for the dual-vehicle system is: ; The lateral dynamic safety margin is: ; The longitudinal dynamic safety margin is: ; in, i Represents a self-driving vehicle. T i For the vehicle's kinetic energy, U i For the system potential energy, It is a vehicle quality For vehicles i The speed of the car, For vehicles j The speed of the car, It is the initial time. It is the end time. It is the longitudinal constraint resistance of traffic rules on drivers. It is the virtual driving force generated by the driver driving the target-driven intelligent vehicle. For the driver's longitudinal target driving force, For the driver's lateral target driving force, For vehicles i longitudinal velocity, For vehicles i lateral velocity, and These represent the lateral constraint forces generated by the two lane lines of the lane in which the target vehicle is traveling. Representative vehicle For vehicles The resulting interactive risk force; It is a vehicle and j Following distance in the longitudinal direction, For vehicles and Lateral following distance, It is the lateral safety margin. It is a positive correlation function. Represents the backward time step. For vehicles j The longitudinal velocity.
5. The method according to claim 1, characterized in that, The risk-sensitive sequential decision-making strategy includes: The vehicle driving strategy is optimized and adjusted continuously within the time window to obtain the optimal dynamic objective function expressed in the rolling time domain. The extreme values of the functional are solved using a preset variational method, and the risk-sensitive sequential decision-making strategy is obtained based on the solution results.
6. The method according to claim 5, characterized in that, The optimal dynamic objective function is expressed in the rolling time domain as follows: ; ; in, S For actual action quantity, k Indicates time, It is a cost function. It is the input vector. It is a state vector. It is the target set; For the ideal action, Represents the rolling time domain. For time increments, For a moment To the future The control input value, For a moment To the future The predicted value, This is a penalty item at the end of the process.
7. A risk-sensitive sequential behavior decision-making device for intelligent vehicles, characterized in that, include: The traffic participant driving information acquisition module is used to acquire the driving status information of traffic participants under a preset traffic environment; The dynamic objective function construction module is used to construct a dynamic objective function based on the driving state information and the driver's risk sensitivity. The intelligent vehicle single-walk behavior decision module is used to determine the vehicle single-walk behavior decision strategy based on the dynamic objective function, longitudinal dynamic safety margin, and lateral dynamic safety margin. The behavioral decision cost calculation and strategy selection module is used to identify the driving intentions of surrounding vehicles based on the single-walk behavioral decision strategy, calculate the cost of different behavioral decision strategies adopted by the current vehicle according to the driving intentions, and match the optimal strategy of the current vehicle according to the cost. as well as The risk-sensitive sequential decision-making model construction and optimization module is used to determine the current vehicle's action at the current moment based on the optimal strategy of the current vehicle, and to determine whether the effect of the single-step output behavior decision strategy is consistent with the actual effect of the current vehicle's action output at the current moment based on rolling time domain optimization. If they are inconsistent, the module reacquires the driving status information of traffic participants in the preset traffic environment until the effect of the risk-sensitive sequential decision-making strategy output is consistent with the actual effect of the current vehicle's action output at the current moment, and outputs the optimal trajectory according to the risk-sensitive sequential decision-making strategy.
8. The apparatus according to claim 7, characterized in that, The dynamic objective function construction module is specifically used for: Construct the minimum action quantity in the real physical system, and determine the unified driving objective in the driver's active decision-making process based on the driving state information; Based on the unified driving objective, output a dynamic objective function based on the minimum action amount. The dynamic objective function is: ; in, i For intelligent vehicles, For intelligent vehicles i The dynamic objective function in the decision-making and planning process. It is the initial time. It is the end time. For the Lagrange equations of a two-vehicle system, T i For the vehicle's kinetic energy, U i This represents the system's potential energy.
9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the intelligent vehicle risk-sensitive sequential behavior decision-making method as described in any one of claims 1-6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the intelligent vehicle risk-sensitive sequential behavior decision-making method as described in any one of claims 1-6.
Citation Information
Patent Citations
Automatic driving integrated decision-making method and device, vehicle and storage medium
CN115534998A
Vehicle anthropomorphic decision control method and device, vehicle and storage medium
CN115923833A