Expert prior guidance-based reinforcement learning automatic driving decision-making method

By adopting reinforcement learning methods based on expert prior guidance in autonomous vehicles, building an ‘action-expected’ observation sequence and performing decision-making updates on hidden Markov model, the problem of insufficient interaction capabilities of autonomous vehicles in unclear right of way scenarios is solved, and higher driving safety and efficiency are achieved.

CN120096625APending Publication Date: 2025-06-06TONGJI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510380315.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When faced with scenarios where there is unclear right of way, existing autonomous vehicles lack initiative and adaptability, resulting in insufficient interaction capabilities and prone to misjudgment and "deadlock".

Method used

The reinforcement learning autonomous driving decision-making method based on expert prior guidance is adopted, and the interactive scenarios with unclear right of way is identified by obtaining real-time traffic environment information, and the reinforcement learning model guided by human driver experts is used to build an "action-expected" observation sequence, and the decision is updated based on the hidden Markov model, and the optimal decision is selected as the final decision of the autonomous vehicle.

Benefits of technology

It improves the driving safety and efficiency of autonomous vehicles in complex interactive scenarios, enhances their interaction success rate when the right of way is unclear, and reduces the occurrence of misjudgment and "deadlock".

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120096625A_ABST
    Figure CN120096625A_ABST
Patent Text Reader

Abstract

The invention relates to a reinforcement learning automatic driving decision-making method based on expert prior guidance, and the method comprises the steps: obtaining real-time traffic environment information, and recognizing an interaction scene with an indefinite passing right through a passing right judgment model considering the time-varying interaction intention and state; determining a corresponding state according to the interaction scene with the unclear right of pass, and obtaining a multi-modal decision result by using a reinforcement learning model demonstrated and guided by a human driver expert so as to construct an action-expectation observation sequence of the human driver; the observation sequence is processed based on a human driver decision updating module, the time sequence influence between the states before and after the interaction process is calculated, and the optimal decision is selected as the final decision of the automatic driving automobile. Compared with the prior art, the active interaction behavior can be generated for the self-driving automobile, the behavior acceptance of the self-driving automobile is greatly improved, the self-driving automobile can better cope with the scene with the undefined passing right, and the misjudgment and deadlock conditions are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a reinforcement learning autonomous driving decision-making method based on expert prior guidance. Background Art

[0002] As the manufacturing and application of autonomous vehicles accelerate, a large number of vehicles will enter the existing traffic ecosystem, which will inevitably lead to a long-term coexistence with human-driven vehicles. This mixture of autonomous driving and human driving exacerbates the inherent complexity of traffic flow dynamics, posing major challenges to the interaction and adaptability of autonomous vehicles in complex environments. Studies have found that about 31% of autonomous driving disengagements are caused by misjudgment of other road users' behavior predictions, failure to provide expected courtesy, or incorrect responses. This is particularly prominent in scenarios with unclear traffic rights of way (such as intersections without signal control), highlighting the current shortcomings of autonomous driving vehicles in driving capabilities and adaptability under ambiguous traffic priorities, and the need to develop more complete autonomous driving intelligent models.

[0003] In the context of the dynamic evolution of ambiguous traffic priority scenarios, the interaction between human-driven vehicles can be conceptualized as an ongoing communication and decision-making process based on widely accepted social norms. These norms describe the behavioral dependencies between human-driven vehicles and help define traffic priorities. In contrast, due to the lack of understanding and application of these universal social norms, autonomous vehicles adopt a predominantly "defensive" strategy. This approach treats human-driven vehicles as moving obstacles to be avoided, rather than taking proactive interactive actions, such as continuously changing speed or trajectory to detect the human-driven vehicle's yielding reaction, or implicitly indicating its expected behavior.

[0004] When dealing with ambiguity in driving scenarios, experienced drivers tend to cope with uncertainty through effective interactions. Rather than passively accepting probabilities, human experts adopt proactive conscious behaviors, including speed changes and trajectory adjustments, to reduce entropy and promote convergence of interactions. Although these decisions may not be optimal, they are simple, effective, and proactive strategies, especially in complex interactive environments.

[0005] Therefore, it is crucial for autonomous vehicles to learn the social norms of interaction between human-driven cars to improve their ability to integrate into existing traffic scenarios, especially when dealing with perceived "deadlock" situations. Summary of the invention

[0006] The purpose of the present invention is to overcome the defect of the traditional autonomous driving decision-making method in the above-mentioned prior art that the interaction with human-driven vehicles when facing unclear right of way lacks initiative, and to provide an autonomous driving decision-making method based on reinforcement learning guided by experts.

[0007] The purpose of the present invention can be achieved by the following technical solutions:

[0008] A reinforcement learning autonomous driving decision-making method based on expert prior guidance includes the following steps:

[0009] Obtain real-time traffic environment information and use a pre-built right-of-way judgment model that considers time-varying interaction intentions and states to identify interaction scenarios where the right of way is unclear;

[0010] Determine the corresponding state according to the interactive scenario with unclear right of way, and use the pre-built reinforcement learning model guided by human driver expert demonstration to obtain multi-modal decision results to construct the action-expected observation sequence of the human driver;

[0011] Based on the pre-built human driver decision update module, the observation sequence is processed, the temporal impact between the states before and after the interaction process is calculated, and the optimal decision is selected as the final decision of the autonomous driving car.

[0012] Furthermore, the right-of-way judgment model considering time-varying interaction intention and state judges whether it is an interaction scenario with unclear right of way based on the expected acceleration, and if the acceleration of the vehicle during the interaction is greater than the expected acceleration, it is an interaction scenario with unclear right of way;

[0013] The calculation expression of the expected acceleration is:

[0014] a d =2(d l -v l T s +βd min ) / T s 2

[0015] In the formula, a d is the expected acceleration, T s represents the time when the interacting vehicles arrive at the conflict point, d l represents the distance between the vehicle and the conflict point, v l represents the speed of the vehicle, d min Indicates the minimum expected gap. β is -1 when calculating sprint acceleration and 1 when calculating yield acceleration.

[0016] Further, the human driver decision updating module performs decision updating based on a hidden Markov model;

[0017] The training process of the human driver decision update module includes:

[0018] Based on the current state of the autonomous driving vehicle and the corresponding multimodal decision results obtained by the reinforcement learning model, as well as the state characteristics of the interactive object vehicle and the corresponding decision results, an action-expectation observation sequence of the human driver is constructed; based on the observation sequence, the hidden variables of the hidden Markov model are estimated, and then combined with the observation sequence to obtain a hidden state sequence;

[0019] The joint distribution of the observation sequence and the hidden state sequence is calculated to obtain the expectation based on conditional probability, which is used to iteratively optimize the parameters of the hidden Markov model until convergence.

[0020] Furthermore, the expected calculation expression based on conditional probability is:

[0021]

[0022] In the formula, is the expectation based on conditional probability, is the current parameter of the hidden Markov model, θ is the model parameter of the previous iteration, D is the total number of observation sequences, I is the hidden state sequence, For a given observation sequence O and parameter The posterior probability of the hidden state sequence I under the condition of , P(O,I|θ) is the joint probability of the observation sequence O and the hidden state sequence I appearing together under the condition of a given parameter θ;

[0023] The parameter optimization expression of the hidden Markov model is:

[0024]

[0025] Furthermore, the process of the human driver decision updating module performing decision updating is specifically as follows:

[0026] Sampling the multimodal decision results obtained by the reinforcement learning model, combining each sampled action with the state to obtain a hidden state sequence, inputting the hidden state sequence into the trained hidden Markov model, calculating the likelihood probability of each hidden state sequence, and thus selecting the action corresponding to the hidden state sequence with the maximum likelihood probability as the final decision result of the autonomous driving vehicle;

[0027] The likelihood probability α of the hidden state sequence at time T t+1 The calculation expression of (i) is:

[0028]

[0029] Where b i (o t ) represents the probability of observing the state generated by the hidden state at the current time, o t is the observed state at time t, ot+1 is the observed state at time t+1, N represents the number of hidden states, α t (j) is the joint probability that the system is in hidden state j at time t and generates the observation sequence before time t, a ji is the transition probability from hidden state j to hidden state i;

[0030] The calculation expression of the likelihood probability P(O|θ) of the hidden state sequence is:

[0031]

[0032] In the formula, α T (i) is the observation sequence O at time T 1:T =(o 1 ,o 2 ,…,o T ), and the joint probability that the system is in hidden state i;

[0033] The final decision result of the autonomous driving car is obtained by:

[0034]

[0035] In the formula, is the action corresponding to the hidden state sequence with the maximum likelihood probability, P i (O|θ) is the likelihood probability that the multimodal decision results constitute a hidden state sequence, and n is the total number of multimodal decision sets.

[0036] Furthermore, the training process of the reinforcement learning model guided by human driver expert demonstration includes the following steps:

[0037] The obtained environment state and corresponding expert trajectory for training are used to initialize the policy parameters of the agent and the parameters of the discriminator, wherein the discriminator is used to distinguish the state-action pairs from the policy of the agent and the state-action pairs from the expert trajectory;

[0038] The following steps are performed in each training round:

[0039] For each time frame, the agent interacts with the environment according to the current strategy, performs action correction, and obtains a state-action-next state sample; and calculates the advantage function of each time frame, which is used to estimate the difference between the actual reward and the state value; until each time frame is traversed;

[0040] Calculate the cumulative return based on the reward function for each time frame;

[0041] For each inner iteration of m time frames, sample a mini-batch B of one time step t , and update the strategy parameters;

[0042] For each n time frame, a mini-batch B is sampled, which includes the trajectories from the expert and the trajectories from the agent. t , and update the discriminator D ω Parameter ω;

[0043] The training iterations are repeated in each training round until the average reward converges and a policy for driving decision-making is obtained.

[0044] Furthermore, the parameter update expression of the discriminator is:

[0045]

[0046] In the formula, Generate trajectories τ for the agent i The expected value of is the gradient of the discriminator update, D ω (s,a) is the discriminator function, For the expert trajectory τ E expected value.

[0047] Furthermore, the reward function includes demonstration learning, traffic efficiency, safety, comfort and discriminator reward, and the corresponding calculation expression is:

[0048] R safe = -w·λ safe

[0049] R comfort =1 / (1+w·λ offfset 2 )

[0050] R efficient =w·λ efficient ·v scaled

[0051] R arrival = w·d scaled +λ arrival

[0052] R expert = -log(D ω (s,a))

[0053] In the formula, R safe is the security reward, w is the reward weight, λ safe is the indicator vector of vehicle collision, R comfort is the comfort bonus, λ offset is the normalized lateral and angular deviation relative to the centerline of the current lane in Frenet coordinates, R efficient is the traffic efficiency reward, vscaled is the vehicle speed mapped to [0, 1], λ efficient is the indicator vector of whether the vehicle speed is within the specified speed limit range, R arrival For demonstration learning rewards, d scaled is the distance from the vehicle to the destination mapped to [0, 1], λ arrival is the indicator vector for the vehicle to obtain the right of way, R expert is the reward for the discriminator.

[0054] Further, the intelligent agent of the reinforcement learning model guided by human driver expert demonstration includes an Actor-Critic network, which includes an actor network and a Critic network;

[0055] The actor network is used to generate an action distribution derived from the current environment state and sample actions from the action distribution; the parameters of the actor network are updated by gradient ascent to maximize the cumulative reward; the loss function of the actor network is expressed as:

[0056]

[0057] In the formula, is the loss function of the actor network, is the expected calculation for all samples in the batch, is the probability ratio of the new and old strategies is the advantage function estimate, is a clipping function that limits the probability ratio to the range of [1-ε,1+ε], where ε is the clipping parameter;

[0058] The Critic network is used to estimate the expected cumulative return based on the current environment state and the action selected by the actor network, and to update the parameters of the Critic network by reducing the difference between the actual cumulative reward and the expected cumulative return estimated by the Critic network; the loss function of the Critic network is expressed as:

[0059]

[0060] In the formula, is the strategy parameter, α is the learning rate, For parameters The gradient of is the probability of taking action a in state s under the new strategy, is the probability of taking action a in state s under the old strategy, A t is the advantage function value, This is an operation to clip the probability ratio within the range of [1-ε,1+ε].

[0061] Furthermore, the state space of the reinforcement learning model guided by the human driver expert demonstration includes the characteristics of the current autonomous driving vehicle and seven characteristics of the interactive object vehicle, and the seven characteristics of the interactive object vehicle are represented by tuples as {x, y, vx, vy, cos h ,sin h ,a d}, where x is the offset to the observed vehicle on the x-axis, y is the offset to the observed vehicle on the y-axis, vx is the speed of the observed vehicle on the x-axis, vy is the speed of the observed vehicle on the y-axis, cos_h is the trigonometric heading of the observed vehicle, sin_h is the trigonometric heading of the observed vehicle, and a d is the predicted expected acceleration of the observed vehicle.

[0062] Compared with the prior art, the present invention has the following advantages:

[0063] (1) The present invention is to deal with the scenario of unclear right of way. First, a right of way judgment method considering time-varying interaction intention and state is used to identify the interaction scenario of unclear right of way. In this scenario, reinforcement learning is implemented according to the environmental state and expert demonstration data to construct the "action-expectation" observation sequence of human drivers. On this basis, a human driver decision update model is established based on the hidden Markov model, thereby considering the temporal influence between the states before and after the interaction process.

[0064] Expert demonstration data is introduced into reinforcement learning and an "actor-critic" structure is used for iterative training to improve the accuracy of the multimodal decision results obtained by reinforcement learning. Further, the constructed "action-expectation" observation sequence is used to make decisions from a timing perspective through a human driver decision update model. In this process, the difference between the self-expectation of the autonomous vehicle and the dynamic response of the HVs is taken into account, avoiding the potential cumulative errors caused by iterative predictions.

[0065] (2) The present invention proposes an active perception decision-making technology that combines expert priors when the right of way is unclear, and uses a parameterized decision update mechanism and a learning strategy to generate active interaction behaviors for autonomous vehicles. It achieves a balance between driving safety and efficiency in complex interaction scenarios, and obtains a higher interaction success rate based on the active interaction strategy. It can optimize the right of way advantage through continuous trial and decision update; it greatly improves the behavioral acceptance of autonomous vehicles, enabling them to better cope with scenarios where the right of way is unclear, and reduce the occurrence of misjudgments and "deadlock" situations.

[0066] (3) In the experimental verification, the present invention constructed a simulation environment that restores the natural driving environment with high precision, and conducted in-depth verification in different driving tasks with unprotected turning scenarios. The results showed that autonomous driving vehicles can accelerate the convergence of interactions through consistent detection and decision updates. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 A flowchart of an autonomous driving decision-making method based on expert prior guidance and reinforcement learning provided by an embodiment of the present invention;

[0068] Figure 2 Calculation of the acceleration threshold required for a turning vehicle to overtake in an embodiment of the present invention;

[0069] Figure 3 Calculation of the acceleration threshold required when a turning vehicle gives way in an embodiment of the present invention;

[0070] Figure 4 For the priority judgment in the initial state of interaction in the embodiment of the present invention, Figure 4 (a) is the initial right of way in the rush-to-pass state; Figure 4 (b) is the initial right of way in the yielding state; Figure 4 (c) in the example indicates that the initial state is unclear;

[0071] Figure 5 It is an HMM-based AV decision update mechanism in an embodiment of the present invention;

[0072] Figure 6 This is the expert demonstration scenario selection in the embodiment of the present invention. The gray area represents the situation where the initial right of way is clear; the scene surrounded by the purple border depicts the ambiguity of the initial right of way, but the turning vehicle passes first, so it can be used as the required expert demonstration data; the scene depicted by the blue border represents the ambiguous initial right of way situation, and the straight-moving vehicle eventually secures the right of way;

[0073] Figure 7 are the training results of the method and the baseline in the embodiment of the present invention; Figure 7 (a) is the average return during training; Figure 7 (b) in the figure is the average travel time during the training process;

[0074] Figure 8 The performance of the method and the baseline in the embodiment of the present invention; Figure 8 (a) is PET; Figure 8 (b) in the equation is the probability of rushing to the front;

[0075] Fig. 9 The intention of obtaining the right of way during the interaction process in the embodiment of the present invention is: Fig. 9(a) in the figure is the baseline strategy (PPO); Fig. 9 (b) in the above is the method proposed in this scheme;

[0076] Fig.10 is the interactive performance (PPO) of the baseline strategy in the embodiment of the present invention;

[0077] Fig.11 is the interactive performance of the proposed strategy in the embodiment of the present invention;

[0078] Fig.12 The interactive performance with HVs under different driving styles in the embodiment of the present invention is: Fig.12 (a) in the figure is the interaction with aggressive HVs; Fig.12 (b) in the figure interacts with the conserved HV;

[0079] Fig.13 is the training performance in the inD dataset in the embodiment of the present invention: Fig.13 (a) in the figure is scenario #2; Fig.13 (b) in the figure is scenario #4;

[0080] Fig.14 This is the interactive performance in the inD dataset in an embodiment of the present invention (scenario #2). DETAILED DESCRIPTION

[0081] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0082] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0083] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0084] Example 1

[0085] like Figure 1 As shown, this embodiment provides a reinforcement learning autonomous driving decision method based on expert prior guidance, including the following steps:

[0086] S1: Obtain real-time traffic environment information and use a pre-built right-of-way judgment model that considers time-varying interaction intentions and states to identify interaction scenarios with unclear right of way;

[0087] S2: Determine the corresponding state according to the interactive scenario with unclear right of way, and use the pre-built reinforcement learning model guided by human driver expert demonstration to obtain multimodal decision results to construct the action-expected observation sequence of the human driver;

[0088] S3: Based on the pre-built human driver decision update module, the observation sequence is processed, the temporal impact between the states before and after the interaction process is calculated, and the optimal decision is selected as the final decision of the autonomous vehicle.

[0089] Specifically, in step S1, the right-of-way judgment model considering the time-varying interaction intention and state is implemented based on the expected acceleration index, which refers to the maximum acceleration required for the vehicle to obtain the right of way or give up the right of way during the interaction. The expression of the expected acceleration is:

[0090] a d =2(d l -v l T s +βd min ) / T s 2

[0091] Where, T s represents the time when the interacting vehicles arrive at the conflict point, d l represents the distance between the vehicle and the conflict point, v l represents the speed of the vehicle, d min Indicates the minimum expected gap, β is -1 when calculating the sprint acceleration and 1 when calculating the yield acceleration. Figure 2 and Figure 3 As shown in FIG. 1 , the calculation process of the expected acceleration threshold when the turning vehicle is trying to overtake and giving way is shown in FIG. 1 .

[0092] like Figure 4 These are the initial right-of-way judgment processes in the rush-to-pass state, yielding state and unclear state.

[0093] In step S2, the expert prior-based reinforcement learning model is determined by the following method:

[0094] The Actor-Critic network in the reinforcement learning model guided by human driver expert demonstrations consists of an actor network for sampling actions based on the agent's current observations, a critic network for evaluating the accumulated reward, and a discriminator network for guiding the agent to approach the expert policy.

[0095] The actor network generates an action distribution derived from the current state of the environment, and then samples actions from this distribution. The parameters of the actor network are updated via gradient ascent, with the goal of maximizing the expected cumulative reward. The loss function of the actor network is:

[0096]

[0097] In the formula, is the loss function of the actor network, is the expected calculation for all samples in the batch, is the probability ratio of the new and old strategies is the advantage function estimate, is a clipping function that limits the probability ratio to the range of [1-ε,1+ε], where ε is the clipping parameter;

[0098] Instead, the Critic network estimates the expected cumulative reward based on the current state of the environment and the actions chosen by the Actor. The parameters of the Critic network are updated by reducing the difference between the actual cumulative reward and the Critic's estimate (i.e., the temporal difference error). The loss function of the Critic network is:

[0099]

[0100] In the formula, is the strategy parameter, α is the learning rate, For parameters The gradient of is the probability of taking action a in state s under the new strategy, is the probability of taking action a in state s under the old strategy, A t is the advantage function value, This is an operation to clip the probability ratio within the range of [1-ε,1+ε].

[0101] The desired policy is obtained based on the discriminator. The discriminator distinguishes whether a given state-action pair originates from the agent's policy or an expert trajectory, with the goal of distinguishing these state-action pair types as much as possible. Through this process, the policy gradually converges to expert behavior.

[0102] A significant advantage of this scheme is that it only requires expert trajectories and no explicit reward function. This benefits it in a variety of situations, including the ones discussed in this study. We characterize the driving behavior of human drivers as a continuous decision-making task, where the expert policy is represented by a sequence of state-action pairs. The parameter update of the discriminator is determined by:

[0103]

[0104] In the formula, for, is the gradient of the discriminator update, and is calculated uniquely for positive and negative trajectories, D ω (s,a) is, for.

[0105] In order to promote consistency between the learned policy and the expert policy, a penalty term can be introduced in the reward function. This modification leads to a revised form of the reward function:

[0106] R expert = -log(D ω (s,a))

[0107] The higher values ​​that the discriminator assigns to state-action pairs indicate that it tends to associate the pairs with expert data.

[0108] The reward function of the reinforcement learning model is determined as follows:

[0109] The reward function motivates the agent to strike a balance between learning from demonstrations, flow efficiency, safety, and comfort. When driving, the agent should strive to complete the unprotected turning task in a stable state without causing collisions (and major conflicts). Specifically, the reward function is set to the following formula:

[0110] R safe = -w·λ safe

[0111] Among them, w represents the reward weight; λ safe An indicator vector indicating vehicle collision (0 or 1);

[0112] R comfort =1 / (1+w·λ offfset 2 )

[0113] Among them, λ offset Represents the normalized lateral and angular deviations relative to the centerline of the current lane in Frenet coordinates, encouraging the vehicle to reduce unnecessary yaw and cornering;

[0114] R efficient =w·λ efficiwnt ·v scaled

[0115] Among them, v scaled represents the vehicle speed mapped to [0, 1], λ efficient Indicates whether the vehicle speed is within the specified speed limit range.

[0116] Rarrival = w·d scaled +λ arrival

[0117] Among them, d scaled represents the distance from the vehicle to the end point mapped to [0, 1]; λ arrival is the indicator vector (0 or 1) that the vehicle has the right of way.

[0118] The observed state space in the Markov process of the expert prior-based reinforcement learning model is determined by Table 1. The state space of the agent includes seven features of the autonomous vehicle and the interacting object vehicle, represented by the tuple {x, y, vx, vy, cos h ,sin h ,a d}, as shown in Table 1.

[0119] Table 1 State space definition

[0120] feature describe x Offset on the x-axis to the observed vehicle. y Offset on the y-axis to the observed vehicle. vx The speed of the observed vehicle on the x-axis. vy The speed of the observed vehicle on the y-axis. cos_h The triangulated heading of the observed vehicle. sin_h The triangulated heading of the observed vehicle. <![CDATA[a d ]]> Predicted expected acceleration of the observed vehicle.

[0121] The training process of the expert prior-based reinforcement learning model includes the following steps:

[0122] The obtained environment state and corresponding expert trajectory for training are used to initialize the agent's strategy parameters and the discriminator's parameters. The discriminator is used to distinguish the state-action pairs from the agent's strategy and the state-action pairs from the expert trajectory.

[0123] The following steps are performed in each training round:

[0124] For each time frame, the agent interacts with the environment according to the current strategy, performs action correction, and obtains a state-action-next state sample; and calculates the advantage function of each time frame, which is used to estimate the difference between the actual reward and the state value; until each time frame is traversed;

[0125] Calculate the cumulative return based on the reward function for each time frame;

[0126] For each inner iteration of m time frames, sample a mini-batch B of one time step t , and update the strategy parameters;

[0127] For each n time frame, a mini-batch B is sampled, which includes the trajectories from the expert and the trajectories from the agent. t , and update the discriminator D ω Parameter ω;

[0128] The training iterations are repeated in each training round until the average reward converges and a policy for driving decision-making is obtained.

[0129] This is equivalent to the algorithm first initializing the policy parameters (θ) and discriminator parameters (ω). The policy is represented by πθ, which determines the actions of the agent in the environment. The discriminator is represented by D ω , attempting to distinguish between state-action pairs coming from the agent’s policy and those coming from the expert’s trajectories.

[0130] In each iteration: (1) the agent interacts with the environment according to the current policy πθ, corrects the action according to the adjustment mechanism, and collects state-action-next-state samples (s, a, s'). (2) At the same time, the algorithm calculates the cumulative reward R for each time step t This is the sum of the rewards from the current time step to the end of the episode, with a discount factor applied to future rewards. (3) In addition, the algorithm calculates the advantage function A at each time step t This involves estimating the difference between the actual reward and the state value, indicating how much better an action is than the average action for that state. (4) In the m inner iterations, a mini-batch B of one time step is sampled. t , and update the policy parameters θ. (5) In the n inner iterations, sample a mini-batch B including correct (from the expert trajectory) and wrong (from the agent trajectory) samples t , and update the discriminator D ω Parameter ω. Used to calculate the updated log-likelihood ratio l, which has different signs for correct and incorrect trajectories. The above process is repeated until the average reward converges. Finally, the policy π that can be used for driving decision making is obtained. θ .

[0131] like Figure 6 The figure shows the selection of expert demonstration scenarios. The gray area indicates the situation where the initial right of way is clear. The scene surrounded by the purple border depicts the ambiguity of the initial right of way, but the turning vehicle passes first, so it can be used as the required expert demonstration data. The scene depicted by the blue border represents an ambiguous initial right of way situation, and the straight-moving vehicle eventually secures the right of way.

[0132] In step S3, the human driver decision update module performs decision update based on the hidden Markov model G-HMM;

[0133] like Figure 5 As shown in Figure 2, the training process of the human driver decision update module includes:

[0134] Based on the current state of the autonomous driving vehicle and the corresponding multimodal decision results obtained by the reinforcement learning model, as well as the state characteristics of the interactive object vehicle and the corresponding decision results, the action-expectation observation sequence of the human driver is constructed; based on the observation sequence, the hidden variables of the hidden Markov model are estimated, and then combined with the observation sequence to obtain the hidden state sequence;

[0135] The joint distribution of the observation sequence and the hidden state sequence is calculated to obtain the expectation based on conditional probability, which is used to iteratively optimize the parameters of the hidden Markov model until convergence.

[0136] Equivalently, the human driver decision update module estimates the hidden variables of the hidden Markov model G-HMM based on the human driver's "action-expectation" observation sequence. For each sample, we first calculate the joint distribution of the observation sequence and the hidden state sequence, and obtain the expectation based on conditional probability:

[0137]

[0138] in, Represents the current parameters of the model. t represents the observation sequence at the current time, θ t represents the hidden variables of the model at the current time. p(θ t |θ t+1 ) represents the likelihood probability of the control transition distribution, p(O t |θ t ) represents the distribution of the control observation value by the likelihood probability. p(θ t |O t ) represents the posterior distribution conditional probability, i.e., the potential decision update mechanism of the human driver.

[0139] The parameter optimization process of the human driver decision update module is expressed as follows:

[0140]

[0141] The above formula is iterated continuously until the parameters converge.

[0142] The process of decision updating by the human driver decision update module is as follows:

[0143] The multimodal decision results obtained by the reinforcement learning model are sampled, and each sampled action is combined with the state to obtain a hidden state sequence, which is input into the trained hidden Markov model to calculate the likelihood probability of each hidden state sequence, so as to select the action corresponding to the hidden state sequence with the maximum likelihood probability as the final decision result of the autonomous driving car;

[0144] Equivalently, the human driver decision update module decodes the G-HMM model to update the autonomous driving decision in the following way. Specifically, the policy network of the RL algorithm samples the feasible action set and combines each action in the set with the state to form a hidden state sequence Then the sequence O=(o 1 ,o 2 ,…,o T) is input into the trained distribution to calculate the likelihood probability of each sequence. The likelihood probability α of the sequence at T t+1 (i) can be calculated as:

[0145]

[0146] Among them, b i (o t ) represents the probability of observing the state generated by the hidden state at the current time, and N represents the number of hidden states.

[0147] The probability of the likelihood of the hidden state sequence is:

[0148]

[0149] Likelihood probability means the prior difference between the action-expectation sequence, so the action with the highest probability is selected - which is also the action with the smallest decision bias. The final action at this moment can be described as:

[0150]

[0151] Both the RL agent and natural driving data exist in a continuous space, and the agent decision needs to be updated based on the state transition probability of each node.

[0152] like Figure 7 As shown in Figure 1, the average reward and average travel time of the training results of the autonomous driving decision-making method proposed in this scheme and the baseline are compared; Figure 8 As shown in Figure 1, the PET and preemption probability comparison of the training results of the autonomous driving decision method proposed in this solution and the baseline is shown in Figure 2. Fig. 9 As shown in Figure 1, it is a comparison of the intention of the baseline strategy (PPO) and the autonomous driving decision-making method proposed in this scheme to obtain the right of way during the interaction process; Fig.10 As shown in, it is the interactive performance (PPO) of the baseline method; Fig.11 As shown in, this is the interactive performance of the autonomous driving decision-making method proposed in this scheme; Fig.12 As shown in Figure 2, the interaction performance of this solution with HVs under different driving styles is shown in Figure 2. Fig.13 As shown in, this is the training performance of this scheme in the inD dataset, Fig.14 As shown, the interactive performance of this scheme in the inD dataset (scenario #2).

[0153] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A reinforcement learning autonomous driving decision-making method based on expert prior guidance, characterized in that: The following steps are involved: Obtain real-time traffic environment information and use a pre-built right-of-way judgment model that considers time-varying interaction intentions and states to identify interaction scenarios where the right of way is unclear; Determine the corresponding state according to the interactive scenario with unclear right of way, and use the pre-built reinforcement learning model guided by human driver expert demonstration to obtain multi-modal decision results to construct the action-expected observation sequence of the human driver; Based on the pre-built human driver decision update module, the observation sequence is processed, the temporal impact between the states before and after the interaction process is calculated, and the optimal decision is selected as the final decision of the autonomous driving car.

2. The method for autonomous driving decision-making based on expert prior guidance through reinforcement learning according to claim 1, characterized in that: The right-of-way judgment model considering time-varying interaction intention and state judges whether it is an interaction scenario with unclear right of way based on expected acceleration, and if the acceleration of the vehicle during interaction is greater than the expected acceleration, it is an interaction scenario with unclear right of way; The calculation expression of the expected acceleration is: a d =2(d l -v l T s +βd min ) / T s 2 In the formula, a d is the expected acceleration, T s represents the time when the interacting vehicles arrive at the conflict point, d l represents the distance between the vehicle and the conflict point, v l represents the speed of the vehicle, d min Indicates the minimum expected gap. β is -1 when calculating sprint acceleration and 1 when calculating yield acceleration.

3. The method for autonomous driving decision-making based on expert prior guidance and reinforcement learning according to claim 1, characterized in that: The human driver decision update module performs decision updates based on a hidden Markov model; The training process of the human driver decision update module includes: Based on the current state of the autonomous driving vehicle and the corresponding multimodal decision results obtained by the reinforcement learning model, as well as the state characteristics of the interactive object vehicle and the corresponding decision results, an action-expectation observation sequence of the human driver is constructed; based on the observation sequence, the hidden variables of the hidden Markov model are estimated, and then combined with the observation sequence to obtain a hidden state sequence; The joint distribution of the observation sequence and the hidden state sequence is calculated to obtain the expectation based on conditional probability, which is used to iteratively optimize the parameters of the hidden Markov model until convergence.

4. The method for autonomous driving decision-making based on expert prior guidance through reinforcement learning according to claim 3, characterized in that: The expected calculation expression based on conditional probability is: In the formula, is the expectation based on conditional probability, is the current parameter of the hidden Markov model, θ is the model parameter of the previous iteration, D is the total number of observation sequences, I is the hidden state sequence, For a given observation sequence O and parameter The posterior probability of the hidden state sequence I under the condition of , P(O,I|θ) is the joint probability of the observation sequence O and the hidden state sequence I appearing together under the condition of a given parameter θ; The parameter optimization expression of the hidden Markov model is:

5. The method for autonomous driving decision-making based on reinforcement learning guided by expert priori according to claim 3, characterized in that: The process of the human driver decision updating module performing decision updating is specifically as follows: Sampling the multimodal decision results obtained by the reinforcement learning model, combining each sampled action with the state to obtain a hidden state sequence, inputting the hidden state sequence into the trained hidden Markov model, calculating the likelihood probability of each hidden state sequence, and thus selecting the action corresponding to the hidden state sequence with the maximum likelihood probability as the final decision result of the autonomous driving vehicle; The likelihood probability α of the hidden state sequence at time T t+1 The calculation expression of (i) is: Where b i (o t ) represents the probability of observing the state generated by the hidden state at the current time, o t is the observed state at time t, o t+1 is the observed state at time t+1, N represents the number of hidden states, α t (j) is the joint probability that the system is in hidden state j at time t and generates the observation sequence before time t, a ji is the transition probability from hidden state j to hidden state i; The calculation expression of the likelihood probability P(O|θ) of the hidden state sequence is: In the formula, α T (i) is the observation sequence O at time T 1:T =(o1,o2,…,o T ), and the joint probability that the system is in hidden state i; The final decision result of the autonomous driving car is obtained by: In the formula, is the action corresponding to the hidden state sequence with the maximum likelihood probability, P i (O|θ) is the likelihood probability that the multimodal decision results constitute a hidden state sequence, and n is the total number of multimodal decision sets.

6. The method for autonomous driving decision-making based on expert prior guidance and reinforcement learning according to claim 1, characterized in that: The training process of the human driver expert demonstration guided reinforcement learning model includes the following steps: The obtained environment state and corresponding expert trajectory for training are used to initialize the policy parameters of the agent and the parameters of the discriminator, wherein the discriminator is used to distinguish the state-action pairs from the policy of the agent and the state-action pairs from the expert trajectory; The following steps are performed in each training round: For each time frame, the agent interacts with the environment according to the current strategy, performs action correction, and obtains a state-action-next state sample; and calculates the advantage function of each time frame, which is used to estimate the difference between the actual reward and the state value; until each time frame is traversed; Calculate the cumulative return based on the reward function for each time frame; For each inner iteration of m time frames, sample a mini-batch B of one time step t , and update the strategy parameters; For each n time frame, a mini-batch B is sampled, which includes the trajectories from the expert and the trajectories from the agent. t , and update the discriminator D ω Parameter ω; The training iterations are repeated in each training round until the average reward converges and a policy for driving decision-making is obtained.

7. The method for autonomous driving decision-making based on reinforcement learning guided by expert priori according to claim 6, characterized in that: The parameter update expression of the discriminator is: In the formula, Generate trajectories τ for the agent i The expected value of is the gradient of the discriminator update, D ω (s,a) is the discriminator function, For the expert trajectory τ E expected value.

8. The method for autonomous driving decision-making based on expert prior guidance through reinforcement learning according to claim 7, characterized in that: The reward function includes demonstration learning, traffic efficiency, safety, comfort and discriminator reward, and the corresponding calculation expression is: R safe =-w·λ safe R comfort =1 / (1+w·λ offset 2 ) R efficient =w·λ efficient ·v scaled R arrival =w·d scaled +λ arrival R expert =-log(D ω (s,a)) In the formula, R safe is the security reward, w is the reward weight, λ safe is the indicator vector of vehicle collision, R comfort is the comfort bonus, λ offset is the normalized lateral and angular deviation relative to the centerline of the current lane in Frenet coordinates, R efficient is the traffic efficiency reward, v scaled is the vehicle speed mapped to [0, 1], λ efficient is the indicator vector of whether the vehicle speed is within the specified speed limit range, R arrival For demonstration learning rewards, d scaled is the distance from the vehicle to the destination mapped to [0, 1], λ arrival is the instruction vector for the vehicle to obtain the right of way, R expert is the reward for the discriminator.

9. The method for autonomous driving decision-making based on expert prior guidance and reinforcement learning according to claim 1, characterized in that: The intelligent agent of the reinforcement learning model guided by human driver expert demonstration includes an Actor-Critic network, which includes an actor network and a Critic network; The actor network is used to generate an action distribution derived from the current environment state and sample actions from the action distribution; the parameters of the actor network are updated by gradient ascent to maximize the cumulative reward; the loss function of the actor network is expressed as: In the formula, is the loss function of the actor network, is the expected calculation for all samples in the batch, is the probability ratio of the new and old strategies is the advantage function estimate, is a clipping function that limits the probability ratio to the range of [1-ε,1+ε], where ε is the clipping parameter; The Critic network is used to estimate the expected cumulative return based on the current environment state and the action selected by the actor network, and to update the parameters of the Critic network by reducing the difference between the actual cumulative reward and the expected cumulative return estimated by the Critic network; the loss function of the Critic network is expressed as: In the formula, is the strategy parameter, α is the learning rate, For parameters The gradient of is the probability of taking action a in state s under the new strategy, is the probability of taking action a in state s under the old strategy, A t is the advantage function value, This is an operation to clip the probability ratio within the range of [1-ε,1+ε].

10. The method for autonomous driving decision-making based on reinforcement learning guided by expert priori according to claim 1, characterized in that: The state space of the reinforcement learning model guided by human driver expert demonstration includes the features of the current autonomous driving vehicle and seven features of the interactive object vehicle. The seven features of the interactive object vehicle are represented by tuples as {x, y, vx, vy, cos h ,sin h ,a d }, where x is the offset to the observed vehicle on the x-axis, y is the offset to the observed vehicle on the y-axis, vx is the speed of the observed vehicle on the x-axis, vy is the speed of the observed vehicle on the y-axis, cos_h is the trigonometric heading of the observed vehicle, sin_h is the trigonometric heading of the observed vehicle, and a d is the predicted expected acceleration of the observed vehicle.

Citation Information

Cited By

  • Taking-over request intensity dynamic adjustment method for aligning cognitive state of driver

    CN121777977A

  • A method for dynamically adjusting the strength of a takeover request in alignment with the cognitive state of the driver

    CN121777977B