An anthropomorphic lane-changing trajectory optimization method based on driver personalization

By combining the basis function with driving style recognition, a maximum entropy inverse reinforcement learning model is constructed, candidate lane change trajectory is generated and optimal trajectory is evaluated and selected, which solves the problem that autonomous vehicles are difficult to achieve anthropomorphic and personalized lane change behavior, and efficient and personalized decisions on anthropomorphic lane change behavior are achieved.

CN116534055BActive Publication Date: 2025-06-17TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310592745.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-24
Publication Date
2025-06-17
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

It is difficult for existing autonomous vehicles to achieve anthropomorphic driving behavior, resulting in low trust in them, and the implementation of personalized lane change behavior is complex and has a large amount of calculations, making it difficult to apply to reality.

Method used

By combining the basis function with driving style recognition, a maximum entropy inverse reinforcement learning model is constructed, candidate lane change trajectory is generated, and the optimal trajectory is selected through reward function evaluation, personalized anthropomorphic lane change behavior decisions are realized.

Benefits of technology

It realizes personalized decisions based on anthropomorphism, reduces the complexity and calculation volume of model, increases the similarity between the driving behavior of autonomous vehicles and humans, and thus increases people's trust in them.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116534055B_ABST
    Figure CN116534055B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for optimizing an anthropomorphic lane-changing trajectory based on driver personalization, including: constructing constraints for the lane-changing trajectory according to the vehicle state at the starting point of the vehicle lane-changing behavior decision, and generating candidate lane-changing trajectories by traversing the vehicle states at the end points of all possible lane-changing behavior decisions; constructing a maximum entropy inverse reinforcement learning model, where the reward function of the maximum entropy inverse reinforcement learning model consists of a set of basis functions, and the basis functions are designed from two aspects of driving safety and personalized driving style to ensure the safety of vehicle lane-changing and personalized driving style. At the same time, the coefficients of the basis functions are determined through inverse reinforcement learning; performing anthropomorphic training on the maximum entropy inverse reinforcement learning model based on the demonstration trajectories of human drivers; evaluating the candidate lane-changing trajectories based on the reward function of the trained maximum entropy inverse reinforcement learning model, and selecting the optimal lane-changing trajectory. Compared with the prior art, the present invention has the advantages of being anthropomorphic, personalized, and having strong safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous vehicle behavior decision-making, and more particularly to an anthropomorphic lane-changing trajectory optimization method based on driver personalization. Background Art

[0002] Autonomous driving has been a hot topic widely studied in academia and industry in recent years. With the development of technology, it can be predicted that there will be more and more autonomous vehicles on the road in the future. However, at the present stage, people's acceptance of autonomous vehicles is not high. A major reason is that the driving behaviors of autonomous vehicles are quite different from those of human drivers, and anthropomorphic driving cannot be achieved. Research shows that the more similar the behaviors of autonomous vehicles are to those of human drivers, the higher the trust people have in autonomous vehicles. Therefore, the driving behaviors of autonomous vehicles should be as similar as possible to those of human drivers, so that autonomous vehicles can be accepted by more people.

[0003] Regarding the simulation of anthropomorphic driving behaviors, with the development of artificial intelligence, there are currently two main methods: imitation learning and inverse reinforcement learning. Imitation learning directly learns the driving behaviors of humans. Inverse reinforcement learning learns the reward function behind human driving behaviors, and then learns the optimal behavior strategy through the reward function. Since the reward function is more transferable in nature compared to directly learning specific behaviors, the generalization ability of inverse reinforcement learning is usually stronger. Usually, inverse reinforcement learning is combined with reinforcement learning. The reward function is obtained through inverse reinforcement learning, and then the optimal strategy is found according to this reward function through reinforcement learning. However, this method makes the model relatively complex and the computational amount relatively large, posing challenges to the application of autonomous vehicles.

[0004] At the same time, with the development of social technology, people are increasingly inclined to personalize products and can customize products according to their own needs. In the case of autonomous vehicles, people hope that autonomous vehicles can drive according to their desired driving styles. Driving style is a summary of the driving behaviors of human drivers and reflects the driving behaviors of human drivers. Different drivers have different driving styles, and different people have different desired driving styles. Therefore, autonomous vehicles should be able to drive according to the desired driving styles of people to achieve personalized autonomous driving.

[0005] The driving behaviors of vehicles can be divided into two main categories: car-following behavior and lane-changing behavior. The interaction object of car-following behavior is only the vehicles in front and behind in the same lane, while the interaction object of lane-changing behavior involves the vehicles in the target lane, which is more complex and dangerous. Therefore, it is very challenging for autonomous vehicles to safely and efficiently achieve personalized and anthropomorphic lane-changing behaviors. Summary of the Invention

[0006] The object of the present invention is to provide a method for optimizing anthropomorphic lane-changing trajectories based on driver personalization, which combines basis functions with driving style recognition to achieve personalized decision-making on the basis of anthropomorphism, and at the same time reduces the model complexity and computational amount by generating candidate lane-changing trajectories.

[0007] The object of the present invention can be achieved by the following technical solutions:

[0008] A method for optimizing anthropomorphic lane-changing trajectories based on driver personalization, comprising the following steps:

[0009] S1. According to the vehicle state at the starting point of the vehicle lane-changing behavior decision, construct the constraints of the lane-changing trajectory, and generate candidate lane-changing trajectories by traversing the vehicle states at the end points of all possible lane-changing behavior decisions;

[0010] S2. Construct a maximum entropy inverse reinforcement learning model, the reward function of the maximum entropy inverse reinforcement learning model consists of a set of basis functions, the basis functions are designed from two aspects of driving safety and personalized driving style to ensure the safety of vehicle lane-changing and personalized driving style, and at the same time the coefficients of the basis functions are determined by inverse reinforcement learning;

[0011] S3. Perform anthropomorphic training on the maximum entropy inverse reinforcement learning model based on the demonstration trajectories of human drivers;

[0012] S4. Evaluate the candidate lane-changing trajectories based on the reward function of the trained maximum entropy inverse reinforcement learning model, and select the optimal lane-changing trajectory.

[0013] Further, in step S1, when driving on a straight road, after the lane-changing behavior decision, the vehicle will enter a stable driving state, approximately a uniform linear motion, so the target state space is:

[0014]

[0015] where v e and y e are respectively the longitudinal speed and the lateral position of the vehicle at the end of the lane-changing process.

[0016] Further, the constraints of the lane-changing trajectory are:

[0017]

[0018]

[0019] Among them, there are 5 constraints for longitudinal motion, namely the initial longitudinal state constraint and the longitudinal target state constraint, and there are 6 constraints for lateral motion, namely the initial lateral state constraint and the lateral target state constraint; x s v s as , y s , v ys , a ys are respectively the longitudinal position, longitudinal speed, longitudinal acceleration, lateral position, lateral speed, and lateral acceleration of the vehicle at the start moment of the decision-making process, that is, the initial state constraint; T is the end moment of the lane-changing process; v e , y e are respectively the longitudinal speed and lateral position of the vehicle at the end moment of the lane-changing process, which are the target state constraints to be determined.

[0020] Furthermore, the candidate lane-changing trajectory is represented as a 4th-degree polynomial in the longitudinal direction and a 5th-degree polynomial in the lateral direction according to the constraint equation of the lane-changing trajectory, that is:

[0021]

[0022] where x(t) is the longitudinal position of the vehicle at time t, and y(t) is the lateral position of the vehicle at time t.

[0023] Furthermore, in the maximum entropy inverse reinforcement learning model, the basis function for measuring driving safety is constructed based on three aspects: following safety, interaction safety, and collision penalty:

[0024] f s (t) = (f follow (t), f interaction (t), f collision (t))

[0025] where f s (t) is the basis function for measuring driving safety, f follow (t) is the basis function for ensuring following safety, f interaction (t) is the basis function for ensuring interaction safety, and f collision (t) represents the basis function for the penalty of a collision.

[0026] Furthermore, the basis function for ensuring following safety is constructed based on the improved time to collision MTTC:

[0027]

[0028]

[0029] Δv = v f - v r

[0030] Δa = a f - a r

[0031]

[0032]

[0033] Among them, T th represents the collision time threshold, d represents the relative distance between the front and rear vehicles, and v f , v r respectively represent the speeds of the front and rear vehicles, and a f , a r respectively represent the accelerations of the front and rear vehicles;

[0034] The basic function for ensuring interactive safety is constructed based on the sum of the longitudinal decelerations of the target lane and the following vehicle in the current lane:

[0035] f interaction (t) = -min(a rx (t), 0) - min(a trx (t), 0)

[0036] Among them, a rx (t) and a trx (t) are respectively the accelerations of the following vehicles in the current lane and the target lane;

[0037] The basic function of the penalty for collision occurrence is:

[0038]

[0039] Furthermore, in the maximum entropy inverse reinforcement learning model, the construction process of the basic function for measuring personalized driving style is as follows: design a set of function groups, calculate the function values of each function in the function group based on the demonstration trajectories of human drivers as features, and perform normalization processing. Use the distributed K-means clustering algorithm for feature selection to classify the driving styles of drivers, and select multiple functions corresponding to the features that have the greatest impact on driving style classification from the function group to form the basic function for measuring personalized driving style.

[0040] Furthermore, the functions in the function group include longitudinal motion jerk, lateral motion jerk, longitudinal speed, longitudinal motion acceleration, longitudinal motion deceleration, lateral motion acceleration, lateral motion deceleration, the distance to the nearest leading vehicle in the target lane at the start of lane change, and the distance to the nearest following vehicle in the target lane at the start of lane change.

[0041] Furthermore, the maximum entropy inverse reinforcement learning model uses the gradient descent method to train the coefficients of the basic functions in the reward function, and the objective function is:

[0042]

[0043]

[0044] Among them, τ is the demonstration trajectory of a human driver with the same driving style; is the candidate trajectory; N is the total number of candidate trajectories; r(τ) represents the reward function of the lane-changing trajectory, and f s is a set of basis functions for measuring driving safety, and ω s is the coefficient obtained by corresponding learning, and f p is a set of basis functions for characterizing personalized driving, and ω p is the coefficient obtained by corresponding learning.

[0045] Furthermore, in step S4, for the selection of candidate lane-changing trajectories, the reward functions under different driving styles learned by using maximum entropy inverse reinforcement learning are substituted into the calculation to obtain the rewards of each candidate trajectory under different driving styles. The probability of a candidate lane-changing trajectory being selected under different driving styles is proportional to the exponent of the reward obtained by this trajectory, that is:

[0046]

[0047] Among them, P j (τ) represents the probability of the candidate lane-changing trajectory being selected under driving style j, and r j (τ) represents the reward obtained by this trajectory under driving style j, and N is the total number of candidate trajectories;

[0048] The candidate trajectory with the maximum probability P j (τ) under driving style j is the optimal lane-changing trajectory determined finally under driving style j.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] (1) By combining basis functions with driving style recognition, the present invention realizes personalized decision-making on the basis of anthropomorphism.

[0051] (2) By generating candidate lane-changing trajectories, the present invention realizes the search for the optimal strategy. Compared with methods such as reinforcement learning for searching for the optimal trajectory, the model complexity and computational amount are greatly reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is the flowchart of the method of the present invention.

[0053] Figure 2 is the schematic diagram of generating lane-changing candidate trajectories in an embodiment.

[0054] Figure 3 is the schematic diagram of the construction process of basis functions in an embodiment.

[0055] Figure 4 is the schematic diagram of realizing personalization and anthropomorphism of the reward function in an embodiment. Specific Embodiment

[0056] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manner and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0057] This embodiment discloses a method for simulating anthropomorphic lane-changing behavior based on driver personalization. The reward function under different driving styles is learned through maximum entropy inverse reinforcement learning, and then the lane-changing behavior decision is made through the learned reward function to achieve the simulation of personalized anthropomorphic lane-changing behavior. The proposed method mainly includes two parts: reward function learning and behavior decision-making. In the reward function learning part, the basis functions are divided into two parts: ensuring driving safety and characterizing personalization. Among them, the basis functions for characterizing personalization are obtained through the driving style feature function to achieve the personalization of the reward function. Then, the basis function coefficients are learned through maximum entropy inverse reinforcement learning to achieve the anthropomorphism of the reward function. In the behavior decision-making part, by traversing the target state space, candidate trajectories are generated, and then the trajectories are evaluated through the reward function, and the trajectory with the maximum reward is selected to achieve the lane-changing behavior decision.

[0058] Since the driving scenarios encountered by vehicles are relatively diverse, the driving scenario assumed in this embodiment is a multi-lane straight road with a lane-changing scenario with surrounding vehicle interactions.

[0059] This embodiment is divided into two parts. One is to use maximum entropy inverse reinforcement learning to learn the lane-changing behavior reward function under different driving styles; the other is lane-changing behavior decision-making to achieve personalized anthropomorphic lane-changing behavior.

[0060] For the decision-making of lane-changing behavior, the general steps for a human driver to make a lane-changing behavior decision are: 1. Generate candidate lane-changing trajectories according to the road conditions, 2. Evaluate the rationality of each generated lane-changing trajectory, and 3. Select a lane-changing trajectory to perform the lane-changing behavior. Therefore, the present invention will also simulate this decision-making step, and divide the lane-changing behavior into these three parts, namely generation, evaluation, and selection, to achieve human-like driving. Specifically, as Figure 1 shown, it includes the following steps:

[0061] S1. According to the vehicle state at the starting point of the vehicle lane-changing behavior decision, construct the constraints of the lane-changing trajectory, and generate candidate lane-changing trajectories by traversing the vehicle states at the end points of all possible lane-changing behavior decisions.

[0062] For the generation of candidate lane-changing trajectories, in this embodiment, based on the vehicle state at the starting point of the lane-changing behavior decision, by traversing the vehicle states at the ending points of possible lane-changing behavior decisions, candidate lane-changing trajectories are generated. Among them, the vehicle state at the ending point can be subdivided into a longitudinal target state and a lateral target state, that is, decoupling the lane-changing movement of the vehicle into longitudinal movement and lateral movement for analysis. The candidate lane-changing trajectories can ultimately be represented by polynomial curves, such as Figure 2 as shown

[0063] When driving on a straight road, after the lane-changing behavior decision, the vehicle will enter a stable driving state, approximately a uniform linear motion. Therefore, the target state space is:

[0064]

[0065] where v e and y e are respectively the longitudinal speed and lateral position of the vehicle at the end of the lane-changing process

[0066] In this embodiment, it is assumed that when on a straight road, at the end of a lane-changing decision-making process, the vehicle will enter a stable driving state (uniform linear motion). At this time, the vehicle's longitudinal acceleration, lateral speed, and acceleration are all 0. For ease of analysis, in this embodiment, the vehicle's longitudinal motion and lateral motion are decoupled and represented by polynomials with respect to time respectively. There are 5 constraints for the longitudinal motion, namely the initial longitudinal state constraint and the longitudinal target state constraint; while for the lateral motion, there are 6 constraints, namely the initial lateral state constraint and the lateral target state constraint. The specific forms of the constraints are as follows:

[0067]

[0068]

[0069] where x s v s a s y s v ys a ys are respectively the longitudinal position, longitudinal speed, longitudinal acceleration, lateral position, lateral speed, and lateral acceleration of the vehicle at the start of the decision-making process, that is, the initial state constraints; T is the end time of the lane-changing process, which is a constant; v e y e are respectively the longitudinal speed and lateral position of the vehicle at the end of the lane-changing process, which are the undetermined target state constraints

[0070] Since there are 5 constraints in the longitudinal direction and 6 constraints in the lateral direction, the driving trajectory curve can be decoupled into a 4th-degree polynomial in the longitudinal direction and a 5th-degree polynomial in the lateral direction, that is:

[0071]

[0072] Among them, x(t) is the longitudinal position of the vehicle at time t, and y(t) is the lateral position of the vehicle at time t.

[0073] As can be seen from the above, there are only two undetermined target state constraints v e , y e affecting the number of trajectory generations. In this embodiment, y e is discretized into the center line of the lane, which is 1 for a two-lane road and 3 for a three-lane or more road. Determining the number num(v e ) after discretization of v e enables obtaining a target space size of num(v e ) or 2num(v e ). By traversing the target space, candidate lane-changing trajectories can be generated. Therefore, the number of candidate trajectories is affected by num(v e ). The larger num(v e ) is, the more candidate trajectories there are, and the higher the degree of anthropomorphism, but the greater the computational amount.

[0074] S2. Construct a maximum entropy inverse reinforcement learning model

[0075] The reward function of maximum entropy inverse reinforcement learning consists of a set of basis functions, and the basis functions are constructed based on mechanism knowledge. In this embodiment, the basis functions are divided into two major parts. One is the basis function for ensuring driving safety, and the other is the basis function for characterizing personalization.

[0076] (1) Basis functions for ensuring driving safety

[0077] For lane-changing behavior, behaviors threatening driving safety can be divided into two aspects. One is the following behavior before and after lane-changing, and the other is the interaction behavior with the target lane and the vehicle behind in the current lane during lane-changing. Therefore, in this embodiment, the basis functions for ensuring driving safety are divided into basis functions for ensuring following safety and basis functions for ensuring interaction safety. At the same time, a collision penalty basis function is introduced. That is, the basis function for measuring driving safety is expressed as:

[0078] f s (t) = (f follow (t), f interaction (t), f collision (t))

[0079] Among them, f s (t) is the basis function for measuring driving safety, f follow (t) is the basis function for ensuring following safety, f interaction (t) is the basis function for ensuring interaction safety, and f collision (t) represents the basis function for the penalty of a collision.

[0080] (1) Basic function for ensuring following - vehicle safety

[0081] The improved time - to - collision is often used to evaluate the safety of following a vehicle. Compared with the time - to - collision, it takes into account the accelerations of the front and rear vehicles and is defined as:

[0082]

[0083] Δv=v f -v r

[0084] Δa=a f -a r

[0085]

[0086]

[0087] where d represents the relative distance between the front and rear vehicles, v f 、v r represent the speeds of the front and rear vehicles respectively, and a f 、a r represent the accelerations of the front and rear vehicles respectively. The smaller the MTTC, the less safe the following - vehicle situation. Generally, when MTTC>20s, it can be considered safe.

[0088] Based on these characteristics of MTTC, the basic function for ensuring following - vehicle safety is defined as:

[0089]

[0090] where T th represents the time - to - collision threshold. In this embodiment, T th is taken as 20s.

[0091] According to the above formula, the more dangerous the following - vehicle situation is, the larger the value of the basic function for ensuring following - vehicle safety is.

[0092] (2) Basic function for ensuring interaction safety

[0093] The lane - changing behavior of a human driver will affect the vehicles in the target lane and the following vehicle in the current lane. Bad lane - changing behavior will cause the vehicles in the target lane and the following vehicle in the current lane to decelerate suddenly, and even collide with the own vehicle. Therefore, in this embodiment, the sum of the longitudinal decelerations of the vehicles in the target lane and the following vehicle in the current lane is used as the basic function for ensuring interaction safety:

[0094] f interaction (t)=-min(a rx (t),0)-min(a trx (t),0)

[0095] where arx (t), a trx (t) are the accelerations of the following vehicles in the current lane and the target lane respectively.

[0096] (3) Basis function for collision penalty

[0097] Since the generated candidate lane-changing trajectories may collide with surrounding vehicles, a basis function for collision penalty is introduced to filter out the trajectories that will collide, that is:

[0098]

[0099] (2) Basis functions for characterizing individuality

[0100] Since many features can reflect the driving style of the driver, this embodiment provides a set of function groups (including basis functions for ensuring following safety and interaction safety). Using the lane-changing trajectories demonstrated by human drivers, calculate the values of each function and perform normalization processing as the features for classifying the driving style of the driver. Using the distributed K-means clustering algorithm for feature selection, classify the driving style of the driver and select the three types of features that have the greatest impact on driving style classification (excluding the basis functions for ensuring following safety and interaction safety) as the basis functions for characterizing individual driving.

[0101] This set of functions includes:

[0102] (1) Longitudinal motion jerk. Jerk is the rate of change of acceleration and is an important indicator for measuring ride comfort:

[0103]

[0104] where x(t) is the longitudinal position of the vehicle at time t.

[0105] (2) Lateral motion jerk, which characterizes the change of lateral motion acceleration and the ride comfort of lateral motion:

[0106]

[0107] where y(t) is the lateral position of the vehicle at time t.

[0108] (3) Longitudinal speed, which measures the driving efficiency of the vehicle:

[0109] f(t) = v x (t)

[0110] (4) Longitudinal motion acceleration:

[0111] f(t) = max(a x (t), 0)

[0112] (5) Longitudinal motion deceleration:

[0113] f(t) = |min(a x (t), 0)|

[0114] (6) Lateral motion acceleration:

[0115] f(t) = max(a y (t), 0)

[0116] (7) Lateral motion deceleration:

[0117] f(τ) = |min(a y (t), 0)|

[0118] (8) At the start time of lane change, the distance to the leading vehicle closest to the host vehicle in the target lane:

[0119]

[0120] (9) At the start time of lane change, the distance to the trailing vehicle closest to the host vehicle in the target lane:

[0121] f(τ) = x r

[0122] The construction process of the basis function is as Figure 3 shown.

[0123] During clustering learning, the above functions are all normalized to prevent the influence of a large value range of a certain function on clustering. After driver driving style classification and feature selection through clustering learning, the basis function of the reward function can be determined. The value of each basis function at each moment is normalized to between [0, 1] to avoid the influence on coefficient learning due to different numerical ranges. Using the demonstration data of human drivers with different driving styles, they are trained separately to obtain the reward function corresponding to each driving style.

[0124] For the learning of the reward function, this embodiment is implemented based on maximum entropy inverse reinforcement learning. Maximum entropy inverse reinforcement learning is a learning algorithm based on the combination of mechanism and data. Its mechanism part is reflected by the basis function of the reward function; while the data part is reflected by learning and training to obtain the coefficients of the basis function. In order to achieve personalization in this embodiment, the basis function is combined with the characteristic function of the driving style. The basis function is divided into two main parts. One is the function to ensure driving safety; the other is the function to represent personalized driving, that is, the characteristic function of the driving style. And through the coefficients of the basis function learned from the demonstration trajectory, the anthropomorphicization of the system is realized, as Figure 4 shown. The learned reward function is:

[0125]

[0126] Among them, τ represents the lane-changing trajectory, r(τ) represents the reward of the lane-changing trajectory, and f s is a set of basis functions to ensure driving safety, and ω s is the coefficient obtained by corresponding learning, and f p is a set of basis functions to characterize personalized driving, and ω p is the coefficient obtained by corresponding learning.

[0127] S3. Perform anthropomorphic training on the maximum entropy inverse reinforcement learning model based on the demonstration trajectories of human drivers.

[0128] After determining the basis functions, learn the coefficient vectors of the basis functions from the demonstration lane-changing trajectories of human drivers with different driving styles through maximum entropy inverse reinforcement learning. In this embodiment, the driving styles of human drivers are divided into three categories (conservative, normal, and aggressive), so three groups of coefficient vectors are learned to obtain three reward functions.

[0129] In this embodiment, the gradient descent method is used to train the coefficients of the basis functions in the reward function, and the objective function is:

[0130]

[0131] Among them, τ is the demonstration trajectory of a human driver with the same driving style; is the candidate trajectory; N is the total number of candidate trajectories; r(τ) is the reward function.

[0132] S4. Evaluate the candidate lane-changing trajectories based on the reward function of the trained maximum entropy inverse reinforcement learning model, and select the optimal lane-changing trajectory.

[0133] For the selection of candidate lane-changing trajectories, use the reward functions under different driving styles learned by maximum entropy inverse reinforcement learning, substitute them into the calculation to obtain the rewards of each candidate trajectory under different driving styles, and the probability of a candidate lane-changing trajectory being selected is proportional to the exponent of the reward obtained by this trajectory, that is:

[0134]

[0135] Among them, P j (τ) represents the probability that the candidate lane-changing trajectory is selected under driving style j, and r j (τ) represents the reward obtained by this trajectory under driving style j, and N is the total number of candidate trajectories;

[0136] The candidate trajectory with the maximum probability P j (τ) under driving style j (the trajectory with the maximum reward function value under driving style j) is the optimal lane-changing trajectory finally determined under driving style j.

[0137] In practical applications, first, the driver conducts a demonstration drive. According to the driver's driving behavior, the driving style of the driver is identified. Then, based on the driving style, the corresponding reward function is selected. Next, the candidate trajectory generated is evaluated by this reward function, and the optimal trajectory is selected, thereby realizing the optimization of personalized anthropomorphic lane-changing trajectories.

[0138] By combining the basis function with driving style recognition, the present invention realizes personalized decision-making on the basis of anthropomorphism. At the same time, through the generation of candidate lane-changing trajectories, the search for the optimal strategy is realized. Compared with methods such as reinforcement learning for searching for the optimal trajectory, this greatly reduces the model complexity and computational amount.

[0139] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations according to the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. A method for optimizing anthropomorphic lane-changing trajectories based on driver personalization, characterized in that, It includes the following steps: S1. Construct the constraints of the lane-changing trajectory according to the vehicle state at the starting point of the vehicle lane-changing behavior decision, and generate candidate lane-changing trajectories by traversing the vehicle states at the ending points of all possible lane-changing behavior decisions; S2. Construct a maximum entropy inverse reinforcement learning model. The reward function of the maximum entropy inverse reinforcement learning model consists of a set of basis functions. The basis functions are designed from two aspects: driving safety and personalized driving style, ensuring the safety of vehicle lane-changing and personalized driving style. At the same time, the coefficients of the basis functions are determined by inverse reinforcement learning; S3. Conduct anthropomorphic training on the maximum entropy inverse reinforcement learning model based on the demonstration trajectories of human drivers; S4. Evaluate the candidate lane-changing trajectories based on the reward function of the trained maximum entropy inverse reinforcement learning model, and select the optimal lane-changing trajectory; In the maximum entropy inverse reinforcement learning model, the construction process of the basis function for measuring the personalized driving style is as follows: Design a set of function groups, calculate the function values of each function in the function group based on the demonstration trajectories of human drivers as features, and perform normalization processing. Use the distributed K-means clustering algorithm for feature selection to classify the driving styles of drivers, and select multiple functions corresponding to the features that have the greatest impact on the driving style classification from the function group to form the basis function for measuring the personalized driving style.

2. The method for optimizing anthropomorphic lane-changing trajectories based on driver personalization according to claim 1, characterized in that, In step S1, when driving on a straight road, after the lane-changing behavior decision, the vehicle will enter a stable driving state, approximately a uniform linear motion. Therefore, the target state space is: Among them, v e and y e are respectively the longitudinal speed and the lateral position of the vehicle at the end of the lane-changing process.

3. The method for optimizing anthropomorphic lane-changing trajectories based on driver personalization according to claim 1, characterized in that, The constraints of the lane-changing trajectory are: Among them, there are 5 constraints for longitudinal motion, namely the initial longitudinal state constraint and the longitudinal target state constraint, and there are 6 constraints for lateral motion, namely the initial lateral state constraint and the lateral target state constraint; x s , v s , a s , y s , v ys , a ys are respectively the longitudinal position, longitudinal speed, longitudinal acceleration, lateral position, lateral speed and lateral acceleration of the vehicle at the start moment of the decision-making process, that is, the initial state constraint; T is the end moment of the lane-changing process; v e , y e are respectively the longitudinal speed and lateral position of the vehicle at the end moment of the lane-changing process, which are the target state constraints to be determined.

4. The method for optimizing anthropomorphic lane-changing trajectories based on driver personalization according to claim 3, characterized in that, The candidate lane-changing trajectories are expressed as a fourth-degree polynomial in the longitudinal direction and a fifth-degree polynomial in the transverse direction according to the constraint equation of the lane-changing trajectory, that is: where x(t) is the longitudinal position of the vehicle at time t, and y(t) is the transverse position of the vehicle at time t.

5. The method for optimizing anthropomorphic lane-changing trajectories based on driver personalization according to claim 1, characterized in that, In the maximum entropy inverse reinforcement learning model, the basis function for measuring driving safety is constructed based on considerations of following safety, interaction safety, and collision penalty: f s (t) = (f follow (t), f interaction (t), f collision (t)) Among them, f s (t) is the basis function for measuring driving safety, f follow (t) is the basis function for ensuring following safety, f interaction (t) is the basis function for ensuring interaction safety, f collision (t) represents the basis function for the penalty of a collision.

6. The method for optimizing anthropomorphic lane-changing trajectories based on driver personalization according to claim 5, characterized in that, The basis function for ensuring following safety is constructed based on the improved time to collision MTTC; Among them, T th represents the collision time threshold, d represents the relative distance between the front and rear vehicles, v f and v r respectively represent the speeds of the front and rear vehicles, a f and a r respectively represent the accelerations of the front and rear vehicles; The basis function for ensuring interaction safety is constructed based on the sum of the longitudinal decelerations of the vehicle behind in the target lane and the current lane; f interaction f(t) = -min(a rx (t), 0) - min(a trx (t), 0) where a rx (t) and a trx (t) are the accelerations of the following vehicles in the current lane and the target lane, respectively; The basis function for the penalty of collision occurrence is:

7. The method for optimizing anthropomorphic lane-changing trajectories based on driver personalization according to claim 1, characterized in that,The functions in the function group include longitudinal motion jerk, transverse motion jerk, longitudinal speed, longitudinal motion acceleration, longitudinal motion deceleration, transverse motion acceleration, transverse motion deceleration, the distance to the vehicle in front closest to the vehicle in the target lane at the start of lane-changing, and the distance to the vehicle behind closest to the vehicle in the target lane at the start of lane-changing.

8. A method for optimizing an anthropomorphic lane-changing trajectory based on driver personalization according to claim 1, wherein, The maximum entropy inverse reinforcement learning model uses the gradient descent method to train the coefficients of the basis functions in the reward function, and the objective function is: where τ is the demonstration trajectory of a human driver with the same driving style; is the candidate trajectory; N is the total number of candidate trajectories; r(τ) represents the reward function of the lane-changing trajectory, f s is a set of basis functions for measuring driving safety, ω s is the coefficient obtained by corresponding learning, f p is a set of basis functions for characterizing personalized driving, ω p is the coefficient obtained by corresponding learning.

9. A method for optimizing an anthropomorphic lane-changing trajectory based on driver personalization according to claim 1, wherein, In step S4, for the selection of candidate lane-changing trajectories, use the reward functions under different driving styles learned by the maximum entropy inverse reinforcement learning, substitute them into the calculation to obtain the rewards of each candidate trajectory under different driving styles. The probability of a candidate lane-changing trajectory being selected under different driving styles is proportional to the exponent of the reward obtained by this trajectory, that is: Among them, P j (τ) represents the probability that the candidate lane-changing trajectory is selected under driving style j, and r j (τ) represents the reward obtained by this trajectory under driving style j, and N is the total number of candidate trajectories; Probability P under driving style j j (τ) The candidate trajectory with the largest value is the optimal lane-changing trajectory under driving style j finally determined.

Citation Information

Patent Citations

  • Heavy commercial vehicle anti-collision early warning method comprehensively considering front and rear obstacles

    CN112622886A

  • Intelligent decision-making and local trajectory planning method for autonomous vehicle and decision-making system thereof

    CN113386795A