Automatic driving decision planning method and system considering interaction and terminal equipment
By designing a joint cost function in the autonomous driving system and using the maximum entropy inverse reinforcement learning algorithm to train the joint reward function, the problem of trajectory prediction and planning of autonomous driving in strong interaction scenarios is solved, and more efficient and safe autonomous driving decisions are achieved.
Patent Information
- Application Number
- CN202510048181.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the strongly interactive autonomous driving scenario, it is difficult for the prior art to accurately predict the trajectory of the traffic vehicle and plan a reasonable trajectory of the main vehicle, and the design of the joint reward function is complex and artificial design is difficult.
By designing a joint cost function, combining safety cost, pass cost and comfort cost, and training through maximum entropy inverse reinforcement learning algorithm, the driving trajectory is automatically planned.
It improves the prediction accuracy of traffic vehicle trajectory, simplifies the design of joint reward function, reduces the need for human intervention, and achieves a more reasonable and safe autonomous driving decision planning.
Smart Images

Figure CN120146218A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and specifically to an autonomous driving decision-making and planning method, system, and terminal device that consider interaction. Background Art
[0002] In strong interaction autonomous driving scenarios such as intersections, ramps, and roundabouts, the key to ensuring the safety of autonomous driving vehicles is to accurately predict the trajectories of traffic vehicles and plan a reasonable trajectory for the host vehicle. Currently, the mainstream decision-making and planning framework decouples the module for predicting traffic vehicle trajectories and the module for planning the host vehicle trajectory. In this framework, the prediction module is upstream of the planning module and does not consider the impact of the future trajectory of the host vehicle on the prediction result. However, the future trajectories of the host vehicle and traffic vehicles affect each other, and there is a complex game relationship between them, rather than a simple causal relationship. The decision-making process of human drivers is a strong proof of this view. Human drivers do not follow this sequential decision-making framework of predicting first and then planning. Instead, they deduce the consequences that different driving behaviors may lead to and select the driving behavior that can optimize the situation from all candidate driving behaviors. Here, the optimal situation refers to the joint optimal situation of the driving goals of the host vehicle and other traffic participants, that is, the situation where the joint reward of all traffic participants including the host vehicle is maximized. Inspired by this, some research has proposed an integrated prediction and planning framework that couples prediction and planning. In this framework, the future trajectory of the host vehicle is considered when predicting the trajectory of traffic vehicles, thus effectively improving the prediction accuracy of traffic vehicle trajectories. However, the joint reward function needs to balance the driving goals of the host vehicle and other traffic participants, including safety, traffic efficiency, and comfort, etc., which makes it very difficult to design the joint reward function artificially. Summary of the Invention
[0003] The purpose of the present invention is to provide an autonomous driving decision-making and planning method, system, and terminal device that consider interaction, aiming to automatically design the joint reward function in the autonomous driving decision-making and planning method and plan a reasonable driving trajectory according to the learned joint reward function.
[0004] According to the first aspect of the present invention, to achieve the above object, the present invention provides the following technical solution: An autonomous driving decision-making and planning method that considers interaction, applied to two-vehicle interactive passing, includes the following steps:
[0005] Design a joint cost function based on safety cost, passing cost, and comfort cost, and perform an opposite number process on the joint cost function to obtain a joint reward function;
[0006] Use the maximum entropy inverse reinforcement learning algorithm to train the joint reward function until the joint reward function converges;
[0007] Sample candidate joint trajectories based on the states of the ego vehicle and the interacting vehicle, and calculate the rewards of the candidate joint trajectories of the ego vehicle and the interacting vehicle using the converged joint reward function;
[0008] Select the candidate joint trajectory with the maximum reward as the planning result for output.
[0009] Furthermore, design a joint cost function based on safety cost, passing cost, and comfort cost, and obtain the joint reward function by taking the negative of the joint cost function, as follows:
[0010] (21) Safety cost design: Quantify the safety cost according to the minimum distance between the ego vehicle and the interacting vehicle, as shown in the following formula:
[0011]
[0012] In the formula, d min represents the closest distance between the outer contours of the two vehicles; D max and D min are preset parameters;
[0013] (22) Passing cost design: Obtained by comparing the passing efficiency of scenarios with and without interaction, and the calculation method of the passing cost is as follows:
[0014]
[0015] In the formula, t bias = t passability - t origin ; t passability represents the time for the vehicle to reach the target position according to the trajectory re-planned due to interaction; t origin represents the time for the vehicle to travel to the target position at the original speed and acceleration; T max and T min are preset parameters;
[0016] (23) Comfort cost design:
[0017] The calculation method is as follows:
[0018]
[0019] In the formula, J Δ represents the absolute value of the difference between the newly planned acceleration and the original acceleration; J max and J min are preset parameters;
[0020] (24) Combining formulas (1) to (3), the joint cost function is expressed as:
[0021] C joint = λ1 C safety +λ 2 C passability1 +λ 3 C passability2 +λ 4 C comfort1 +λ 5 C comfort2 (4)
[0022] In the formula, C passability1 and C passability2 respectively represent the passing costs of the self - vehicle and the interacting vehicle; C comfort1 and C comfort2 respectively represent the comfort costs of the self - vehicle and the interacting vehicle; C safety represents the safety cost of the self - vehicle and the interacting vehicle, and the safety cost is shared by the self - vehicle and the interacting vehicle; λ 1 , λ 2 , λ 3 , λ 4 and λ 5 respectively represent the weight coefficients of each cost;
[0023] (25) Take the opposite of the combined cost function as the combined reward function, specifically as follows:
[0024] R = -C joint .
[0025] Furthermore, use the maximum entropy inverse reinforcement learning algorithm to train the combined reward function until the combined reward function converges, specifically as follows:
[0026] (31) Initialize the parameters of the combined reward function;
[0027] (32) Sample the corresponding candidate combined trajectory set according to the expert demonstration trajectories of each pair of human drivers. Each pair of candidate combined trajectories and the corresponding expert demonstration trajectories have the same initial state;
[0028] (33) Use the current combined reward function to calculate the rewards of the expert demonstration trajectories and the candidate combined trajectories;
[0029] (34) Calculate the gradient of the updated parameters and update the parameters of the combined reward function through the gradient ascent algorithm;
[0030] (35) Repeat steps (33) to (34) until the combined reward function converges.
[0031] Furthermore, the candidate combined trajectory set is composed of pairwise pairing of the candidate trajectories of the self - vehicle and the candidate trajectories of the interacting vehicle. Sample the corresponding candidate combined trajectory set, specifically as follows:
[0032] The trajectories of the ego vehicle and the interaction vehicle are represented as the time series of vehicle states τ = [x 1 ,x 2 …,x L ], the vehicle state x = [S, v, a] includes the longitudinal position S, velocity v and acceleration a of the vehicle, then the longitudinal dynamics model of the vehicle is expressed as:
[0033]
[0034] When the initial state of the vehicle is determined, the trajectory of the vehicle can be determined by a given acceleration sequence within the planning field of view. Therefore, a candidate trajectory set can be constructed by sampling candidate acceleration sequences.
[0035] Furthermore, the current joint reward function is used to calculate the rewards of the expert demonstration trajectory and the candidate joint trajectory, as follows:
[0036] (51) According to the maximum entropy inverse reinforcement learning algorithm, the probability of a trajectory being selected is proportional to the natural exponent of its reward:
[0037]
[0038] Where P(τ|θ) represents the probability of trajectory τ being selected; R(τ|θ) represents the reward of trajectory τ; θ represents the parameter of the joint reward function, θ = [λ 1 ,λ 2 ,λ 3 ,λ 4 ,λ 5 ] T ; Z(θ)=∫ D e R(τ|θ) dτ is called the normalization function or partition function; D represents the set of all trajectories;
[0039] In the sampling-based inverse reinforcement learning method, the partition function Z(θ) is:
[0040]
[0041] In the formula, φ represents the set of sampled candidate trajectories;
[0042] Therefore, the probability of trajectory τ being selected can be expressed as:
[0043]
[0044] In the formula, all trajectories in the sampled candidate trajectory set φ have the same initial state as the trajectory τ;
[0045] (52) The likelihood of the expert demonstration trajectory is maximized by adjusting the parameter θ of the joint reward function, as shown in the following formula:
[0046]
[0047] where \(E\) represents the set of expert demonstration trajectories;
[0048] For the convenience of calculation, the above formula is transformed into:
[0049]
[0050] It can be obtained from Equation (9) that the objective function of the sampling-based maximum entropy deep inverse reinforcement learning algorithm is represented in the following form:
[0051]
[0052] Each expert demonstration trajectory corresponds to a set of candidate trajectories \(\varphi\) e , and all the trajectories in the set of candidate trajectories \(\varphi\) e have the same initial state as their corresponding expert trajectory \(\tau\) e .
[0053] Furthermore, calculate the gradient of the update parameter and update the parameter of the joint reward function through the gradient ascent algorithm, as follows:
[0054] The gradient of the update parameter \(\theta\) is:
[0055]
[0056] where \(R(\tau|\theta)\) represents the reward of trajectory \(\tau\); \(\theta\) represents the parameter of the joint reward function, \(\theta = [\lambda\) 1 , \lambda\) 2 , \lambda\) 3 , \lambda\) 4 , \lambda\) 5 T .
[0057] According to the second aspect of the present invention, the present invention provides an autonomous driving decision-making and planning system considering interaction, which is used to implement the above-mentioned autonomous driving decision-making and planning method considering interaction, including:
[0058] A function design module, which is used to design a joint cost function based on safety cost, passing cost and comfort cost, and obtain a joint reward function by taking the opposite of the joint cost function;
[0059] A training module, which is used to train the joint reward function by using the maximum entropy inverse reinforcement learning algorithm until the joint reward function converges;
[0060] A calculation module, which is used to sample candidate joint trajectories according to the states of the host vehicle and the interaction vehicle, and calculate the rewards of the candidate joint trajectories of the host vehicle and the interaction vehicle by using the converged joint reward function;
[0061] A selection output module is configured to select the candidate joint trajectory with the maximum reward as the planning result for output.
[0062] According to a third aspect of the present invention, there is provided a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, the above-described autonomous driving decision-making and planning method considering interactions is adopted.
[0063] According to a fourth aspect of the present invention, there is provided a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the above-described autonomous driving decision-making and planning method considering interactions when executed by a computer processor.
[0064] The present invention at least has the following beneficial effects:
[0065] By constructing an integrated prediction framework, the present invention fully considers the interaction and game relationship between the autonomous driving vehicle and other traffic vehicles, and through the maximum entropy inverse reinforcement learning algorithm, according to the expert demonstration data of human drivers, the joint reward function of the integrated prediction framework is automatically calibrated.
[0066] Of course, it is not necessary for any product implementing the present invention to simultaneously achieve all the above-mentioned advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 It is a flowchart of the method described in the present invention;
[0068] Figure 2 It is a flowchart of the process of training the joint reward function in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.
[0070] Embodiment 1:
[0071] Please refer to Figure 1 - Figure 2 , the present invention provides a technical solution: an autonomous driving decision-making and planning method considering interactions, including the following steps:
[0072] S1. Design a combined cost function based on safety cost, passing cost, and comfort cost, and obtain a combined reward function by taking the opposite of the combined cost function, as follows:
[0073] (S11) Safety cost design: The safety cost is a measure of the safety of the combined trajectories of the ego vehicle and the interacting vehicle. The safety cost is quantified according to the minimum distance between the ego vehicle and the interacting vehicle, as shown in the following formula:
[0074]
[0075] In the formula, d min represents the closest distance between the outer contours of the two vehicles; D max and D min are preset parameters;
[0076] (S12) Passing cost design: The passing cost is used to quantify the impact of the interaction behavior on the passing efficiency. This impact is obtained by comparing the passing efficiencies of the scenarios with and without interaction. The calculation method of the passing cost is as follows:
[0077]
[0078] In the formula, t bias = t passability - t origin ; t passability represents the time for the vehicle to reach the target position according to the re-planned trajectory due to the interaction; t origin represents the time for the vehicle to reach the target position by traveling at the original speed and acceleration; T max and T min are preset parameters;
[0079] (S13) Comfort cost design: The comfort cost refers to the reduction in comfort caused by the interaction;
[0080] The calculation method is as follows:
[0081]
[0082] In the formula, J Δ represents the absolute value of the difference between the newly planned acceleration and the original acceleration; J max and J min are preset parameters;
[0083] (S14) Combining formulas (1) to (3), the combined cost function is expressed as:
[0084] C joint = λ 1 C safety + λ 2 C passability1 + λ3 C passability2 + λ 4 C comfort1 + λ 5 C comfort2 (4)
[0085] In the formula, C passability1 and C passability2 respectively represent the passing costs of the host vehicle and the interacting vehicle; C comfort1 and C comfort2 respectively represent the comfort costs of the host vehicle and the interacting vehicle; C safety represents the safety cost of the host vehicle and the interacting vehicle, and the safety cost is shared by the host vehicle and the interacting vehicle; λ 1 , λ 2 , λ 3 , λ 4 and λ 5 respectively represent the weight coefficients of each cost;
[0086] (S15) Take the opposite of the combined cost function as the combined reward function, specifically as follows:
[0087] R = -C joint .
[0088] S2. Use the maximum entropy inverse reinforcement learning algorithm to train the combined reward function until the combined reward function converges, specifically as follows:
[0089] (S21) Initialize the parameters of the combined reward function;
[0090] (S22) Sample the corresponding candidate combined trajectory set according to the expert demonstration trajectory of each pair of human drivers. Each pair of candidate combined trajectories and the corresponding expert demonstration trajectory have the same initial state;
[0091] Install a driving data acquisition device on the vehicle to collect the expert demonstration trajectory of the human driver;
[0092] The candidate combined trajectory set is composed of pairwise pairing of the candidate trajectories of the host vehicle and the candidate trajectories of the interacting vehicle. Sample the corresponding candidate combined trajectory set, specifically as follows:
[0093] The trajectories of the host vehicle and the interacting vehicle are represented as a time series of vehicle states τ = [x 1 , x 2 …, x L . The vehicle state x = [S, v, a] includes the longitudinal position S, speed v, and acceleration a of the vehicle. Then the vehicle longitudinal dynamics model is represented as:
[0094]
[0095] Once the initial state of the vehicle is determined, the trajectory of the vehicle can be determined by a given acceleration sequence within the planned horizon. Therefore, a set of candidate trajectories can be constructed by sampling candidate acceleration sequences;
[0096] (S23) Calculate the rewards of the expert demonstration trajectory and the candidate joint trajectory using the current joint reward function, as follows:
[0097] (S23.1) According to the maximum entropy inverse reinforcement learning algorithm, the probability of a trajectory being selected is proportional to the natural exponent of its reward:
[0098]
[0099] In the formula, P(τ|θ) represents the probability that the trajectory τ is selected; R(τ|θ) represents the reward of the trajectory τ (expert demonstration trajectory or candidate joint trajectory); θ represents the parameters of the joint reward function, θ = [λ 1 , λ 2 , λ 3 , λ 4 , λ 5 ; T ; Z(θ) = ∫ D e R(τ|θ) dτ is called the normalization function or partition function; D represents the set of all trajectories;
[0100] In the sampling-based inverse reinforcement learning method, the partition function Z(θ) is:
[0101]
[0102] In the formula, φ represents the set of sampled candidate trajectories;
[0103] Therefore, the probability that the trajectory τ is selected can be expressed as:
[0104]
[0105] In the formula, all the trajectories in the sampled candidate trajectory set φ have the same initial state as the trajectory τ;
[0106] (S23.2) Maximize the likelihood of the expert demonstration trajectory by adjusting the parameters θ of the joint reward function, as shown in the following formula:
[0107]
[0108] In the formula, E represents the set of expert demonstration trajectories;
[0109] For ease of calculation, the above formula is transformed into:
[0110]
[0111] It can be obtained from Equation (9) that the objective function of the sampling-based maximum entropy deep inverse reinforcement learning algorithm is expressed as follows:
[0112]
[0113] Each expert demonstration trajectory corresponds to a set of candidate trajectories φ e , the set of candidate trajectories φ e All the trajectories in have the same initial state as their corresponding expert trajectory τ e ;
[0114] (S24) Calculate the gradient of the updated parameter and update the parameter of the joint reward function through the gradient ascent algorithm as follows:
[0115] The gradient of the updated parameter θ is:
[0116]
[0117] where R(τ|θ) represents the reward of trajectory τ (expert demonstration trajectory or candidate joint trajectory); θ represents the parameter of the joint reward function, θ = [λ 1 , λ 2 , λ 3 , λ 4 , λ 5 T ;
[0118] (S25) Repeat steps (S23) to (S24) until the joint reward function converges;
[0119] S3. Sample candidate joint trajectories according to the states of the ego vehicle and the interacting vehicle, and calculate the rewards of the candidate joint trajectories of the ego vehicle and the interacting vehicle using the converged joint reward function, the same as step (S23);
[0120] S4. Select the candidate joint trajectory with the maximum reward as the planning result for output.
[0121] In summary, the present invention constructs an integrated prediction framework, fully considers the interaction game relationship between the autonomous driving vehicle and other traffic vehicles, and automatically calibrates the joint reward function of the integrated prediction framework according to the expert demonstration data of human drivers through the maximum entropy inverse reinforcement learning algorithm.
[0122] Embodiment 2:
[0123] This embodiment provides an autonomous driving decision-making and planning system considering interaction, which is used to implement the autonomous driving decision-making and planning method considering interaction described in Embodiment 1, including:
[0124] A function design module, which is used to design a combined cost function based on safety cost, passing cost, and comfort cost, and obtain a combined reward function by taking the opposite of the combined cost function;
[0125] A training module, which is used to train the combined reward function using the maximum entropy inverse reinforcement learning algorithm until the combined reward function converges;
[0126] A calculation module, which is used to sample candidate combined trajectories according to the states of the ego vehicle and the interacting vehicle, and calculate the rewards of the candidate combined trajectories of the ego vehicle and the interacting vehicle using the converged combined reward function;
[0127] A selection and output module, which is used to select the candidate combined trajectory with the maximum reward as the planning result for output.
[0128] Specifically, the above function design module, training module, calculation module, and selection and output module can be embedded in a computer processing system. The computer calls the above modules according to the above-provided autonomous driving decision-making and planning method considering interactions to complete the task of planning the passing trajectory; the above function design module, training module, calculation module, and selection and output module can perform operations according to the specific steps given by the above autonomous driving decision-making and planning method considering interactions.
[0129] It should be noted that it should be understood that the division of each module of the above system is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; it is also possible that some modules are implemented in the form of software called by processing elements, and some modules are implemented in the form of hardware. For example, the function design module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above signal processing module is called and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.
[0130] For example, the above-mentioned modules can be one or more integrated circuits configured to implement the above methods. For example: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0131] Embodiment 3:
[0132] The present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The computer program stored in the memory is capable of running on the processor. When the processor loads and executes the computer program, the above-mentioned autonomous driving decision-making and planning method considering interaction is adopted.
[0133] It should be noted that the terminal device can be a computer device such as a desktop computer, a laptop computer, or a cloud server. And the terminal device includes but is not limited to a processor and a memory. For example, the terminal device can also include input / output devices, network access devices, and a bus, etc.
[0134] Furthermore, the processor can adopt a Central Processing Unit (CPU). Of course, according to actual usage, other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), off-the-shelf Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be adopted. The general-purpose processor can adopt a microprocessor or any conventional processor, etc. The present application does not make any restrictions on this.
[0135] Embodiment 4:
[0136] The present invention provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the above-mentioned autonomous driving decision-making and planning method considering interaction when executed by a computer processor.
[0137] Among them, the computer program can be stored in a computer-readable medium. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some middleware form, etc. The computer-readable medium includes any entity or device, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the computer-readable medium includes but is not limited to the above components.
[0138] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0139] For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. When an element is referred to as being "assembled on", "installed on", "fixed to" or "disposed on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used herein are only for the purpose of illustration and do not represent the only implementation.
[0140] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
[0141] In the description of this specification, the description with reference to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
Claims
1. An interactive autonomous driving decision-making planning method is applied to two-car interactive meeting, which is characterized by: The specific steps include: A joint cost function is designed based on safety cost, traffic cost and comfort cost, and a joint reward function is obtained by performing inverse number processing on the joint cost function; The joint reward function is trained using the maximum entropy inverse reinforcement learning algorithm until the joint reward function converges; Sample candidate joint trajectories based on the states of the ego vehicle and the interaction vehicle, and use the converged joint reward function to calculate the rewards of the candidate joint trajectories of the ego vehicle and the interaction vehicle; The candidate joint trajectory with the largest reward is selected as the planning result output.
2. The interactive autonomous driving decision-making planning method according to claim 1, characterized in that: The joint cost function is designed based on the safety cost, traffic cost and comfort cost, and the joint reward function is obtained by performing the inverse number processing on the joint cost function, as follows: (21) Safety cost design: The safety cost is quantified according to the minimum distance between the ego vehicle and the interacting vehicle, as shown in the following formula: Where, d min Indicates the shortest distance between the outer contours of the two vehicles; D max and D min is the preset parameter; (22) Traffic cost design: By comparing the traffic efficiency of interactive and non-interactive scenarios, the calculation method of the traffic cost is as follows: Where, t bias =t passability -t origin ;t passability It represents the time it takes for the vehicle to reach the target location according to the trajectory replanned due to the interaction; t origin It indicates the time it takes for the vehicle to reach the target position according to the original speed and acceleration; T max and T min is the preset parameter; (23) Comfortable design: The calculation method is as follows: In the formula, J Δ Indicates the absolute value of the difference between the new planned acceleration and the original acceleration; J max and J min is the preset parameter; (24) Combining formula (1) to formula (3), the joint cost function is expressed as: C joint =λ1C safety +λ2C passability1 +λ3C passability2 +λ4C comfort1 +λ5C comfort2 (4) In the formula, C passability1 and C passability2 Represent the travel costs of the self-vehicle and the interactive vehicle respectively; C comfort1 and C comfort2 Represent the comfort cost of the self-vehicle and the interactive vehicle respectively; C safety represents the safety cost of the ego vehicle and the interactive vehicle, and the safety cost is shared by the ego vehicle and the interactive vehicle; λ1, λ2, λ3, λ4 and λ5 represent the weight coefficients of each cost respectively; (25) The inverse of the joint cost function is used as the joint reward function, as follows: R=-C joint 。 3. The interactive autonomous driving decision-making planning method according to claim 1, characterized in that: The maximum entropy inverse reinforcement learning algorithm is used to train the joint reward function until the joint reward function converges, as follows: (31) Initialize the parameters of the joint reward function; (32) sampling a corresponding candidate joint trajectory set according to each pair of expert demonstration trajectories of human drivers, where each pair of candidate joint trajectories and the corresponding expert demonstration trajectory have the same initial state; (33) Calculate the rewards of the expert demonstration trajectory and the candidate joint trajectory using the current joint reward function; (34) Calculate the gradient of the updated parameters and update the parameters of the joint reward function through the gradient ascent algorithm; (35) Repeat steps (33) to (34) until the joint reward function converges.
4. The interactive autonomous driving decision-making planning method according to claim 3, characterized in that: The candidate joint trajectory set is composed of the candidate trajectories of the ego vehicle and the candidate trajectories of the interactive vehicle in pairs. The corresponding candidate joint trajectory set is sampled as follows: The trajectories of the ego vehicle and the interaction vehicle are represented as a time series of vehicle states τ = [x1, x2…, x L ], the vehicle state x = [S, v, a] includes the longitudinal position S, velocity v and acceleration a of the vehicle, then the longitudinal dynamics model of the vehicle is expressed as: When the initial state of the vehicle is determined, the trajectory of the vehicle can be determined given the acceleration sequence within the planning field of view. Therefore, a candidate trajectory set can be constructed by sampling candidate acceleration sequences.
5. The interactive autonomous driving decision-making planning method according to claim 4, characterized in that: The current joint reward function is used to calculate the rewards of the expert demonstration trajectory and the candidate joint trajectory as follows: (51) According to the maximum entropy inverse reinforcement learning algorithm, the probability of a trajectory being selected is proportional to the natural exponent of its reward: Where P(τ|θ) represents the probability of trajectory τ being selected; R(τ|θ) represents the reward of trajectory τ; θ represents the parameter of the joint reward function, θ = [λ1,λ2,λ3,λ4,λ5] T ; Z(θ)=∫ D e R(τ|θ) dτ is called the normalization function or partition function; D represents the set of all trajectories; In the sampling-based inverse reinforcement learning method, the partition function Z(θ) is: In the formula, φ represents the set of sampled candidate trajectories; Therefore, the probability of trajectory τ being selected can be expressed as: In the formula, all trajectories in the sampled candidate trajectory set φ have the same initial state as the trajectory τ; (52) The likelihood of the expert demonstration trajectory is maximized by adjusting the parameter θ of the joint reward function, as shown in the following formula: Where E represents the set of expert demonstration trajectories; For the convenience of calculation, the above formula is transformed into: From formula (9), the objective function of the sampling-based maximum entropy deep inverse reinforcement learning algorithm is expressed as follows: Each expert demonstration trajectory corresponds to a candidate trajectory set φ e , the candidate trajectory set φ e All trajectories in and their corresponding expert trajectories τ e have the same initial state.
6. The interactive autonomous driving decision-making planning method according to claim 5, characterized in that: Calculate the gradient of the updated parameters and update the parameters of the joint reward function through the gradient ascent algorithm as follows: The gradient of updating parameter θ is: Where R(τ|θ) represents the reward of trajectory τ; θ represents the parameters of the joint reward function, θ = [λ1,λ2,λ3,λ4,λ5] T .
7. An autonomous driving decision-making planning system considering interaction, used to implement the autonomous driving decision-making planning method considering interaction as claimed in any one of claims 1 to 6, characterized in that: include: A function design module is used to design a joint cost function based on safety cost, traffic cost and comfort cost, and to perform inverse number processing on the joint cost function to obtain a joint reward function; A training module, used to train the joint reward function using a maximum entropy inverse reinforcement learning algorithm until the joint reward function converges; A calculation module is used to sample candidate joint trajectories according to the states of the ego vehicle and the interaction vehicle, and calculate the rewards of the candidate joint trajectories of the ego vehicle and the interaction vehicle using the converged joint reward function; The output selection module is used to select the candidate joint trajectory with the largest reward as the planning result output.
8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on a processor. When the processor loads and executes the computer program, the interactive autonomous driving decision-making planning method described in any one of claims 1 to 6 is adopted.
9. A storage medium containing computer executable instructions, characterized in that: The computer executable instructions are used to execute the interactive autonomous driving decision planning method according to any one of claims 1 to 6 when executed by a computer processor.
Citation Information
Cited By
Traffic control method and device, electronic equipment and computer storage medium
CN120877539A