Vehicle control method and system for automatic driving overtaking scene in combination with path planning and driving style, and storage medium

By combining path planning and driving style methods, using the secondary planning algorithm and deep Q learning strategy to optimize overtaking paths and speeds, the problem of fuel economy not being fully considered in the overtaking scenarios in autonomous driving technology is solved, and efficient fuel economy and driving style are achieved.

CN120080846APending Publication Date: 2025-06-03KUNMING UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510194350.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing autonomous driving technology has failed to effectively combine driving style with fuel economy in overtaking scenarios, resulting in fuel economy not being fully considered.

Method used

By obtaining the historical driving characteristic parameters of the vehicle, determining the target driving style type, and optimizing the overtaking path and speed based on the quadratic planning algorithm and deep Q learning strategy to achieve both fuel economy and driving style.

Benefits of technology

It has achieved the reduction of energy consumption under different driving styles, significantly improved fuel economy, and improved the practical application value and popularization potential of autonomous driving technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120080846A_ABST
    Figure CN120080846A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic driving, in particular to a vehicle control method and system for an automatic driving overtaking scene in combination with path planning and a driving style and a storage medium. The method comprises the steps that historical driving characteristic parameters of a vehicle are acquired, and a target driving style type is determined according to the historical driving characteristic parameters; based on a quadratic programming algorithm, determining an overtaking planning path and an overtaking speed meeting the target driving style; formulating an upper-layer speed optimization strategy and an energy management strategy of lower-layer depth Q learning based on the overtaking planned path, the overtaking speed and the target driving style; and controlling the vehicle speed constraint during overtaking of the vehicle to meet the upper-layer speed optimization strategy, and controlling the power distribution between a battery and an engine of the vehicle to meet the energy management strategy of the lower-layer depth Q learning. The invention aims to solve the problem of how to control the vehicle according to the driving style and the fuel economy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and particularly to a vehicle control method, system and storage medium for an autonomous driving overtaking scenario that combines path planning and driving style. Background Art

[0002] A driving assistance system usually needs to detect and analyze the driving condition data of a vehicle and the driving data of a driver to identify the driving style of the driver, so as to provide a more user-friendly service and a safer and more comfortable assisted driving for the driver.

[0003] In related technical solutions, Chinese Patent with application number 202211575828 discloses a vehicle control method, device, electronic device and computer-readable medium. By controlling an on-vehicle sensor to collect the environmental information and interaction information of a target vehicle; identifying the driver information of the target vehicle; in response to determining that the driver information corresponds to a target driver, determining the driving style information of the target driver; according to the environmental information, interaction information and driving style information, identifying the lane-changing intention information of the target driver; according to the lane-changing intention information and environmental information, generating vehicle control information; determining whether there is a driving conflict between the vehicle control information and the interaction information; and controlling the on-vehicle system to perform lane-changing processing on the target vehicle according to the vehicle control information.

[0004] In the above solution, it mainly focuses on generating different lane-changing trajectories according to different driving styles and then controlling the vehicle to change lanes, without considering the fuel economy of the vehicle during autonomous driving.

[0005] The above content is only used to assist in understanding the technical solution of the present application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main purpose of the present application is to provide a vehicle control method for an autonomous driving overtaking scenario that combines path planning and driving style, aiming to solve the problem of how to control a vehicle based on driving style and fuel economy.

[0007] To achieve the above purpose, a vehicle control method for an autonomous driving overtaking scenario that combines path planning and driving style provided by the present application includes:

[0008] Obtain the historical driving characteristic parameters of the vehicle, and determine the target driving style type according to the historical driving characteristic parameters. The historical driving characteristic parameters include the mean lateral acceleration, mean lateral speed, steering wheel angle, standard deviation of lateral speed, standard deviation of lateral acceleration and average acceleration. The target driving styles include an aggressive driving style, a normal driving style and a conservative driving style;

[0009] Based on the quadratic programming algorithm, determine the overtaking planning path and overtaking vehicle speed that meet the target driving style;

[0010] Based on the overtaking planning path, the overtaking vehicle speed, and the target driving style, formulate an upper-layer speed optimization strategy and a lower-layer energy management strategy for deep Q-learning;

[0011] Control the vehicle speed constraint during overtaking to meet the upper-layer speed optimization strategy, and control the power distribution between the vehicle's battery and engine to meet the lower-layer energy management strategy for deep Q-learning.

[0012] Optionally, the step of determining the target driving style type according to the historical driving characteristic parameters includes:

[0013] Reduce the dimension of the historical driving characteristic parameters based on the principal component analysis method to form a preprocessed data set;

[0014] Calculate the Euclidean distance between each point in the preprocessed data set and the center point;

[0015] Perform K-means clustering based on the Euclidean distance, and calculate the center point value of the lateral speed, the center point value of the lateral acceleration, and the center point value of the steering wheel mean;

[0016] Determine the driving style type corresponding to the interval where the center point value of the lateral speed, the center point value of the lateral acceleration, and the center point value of the steering wheel mean are located as the target driving style type.

[0017] Optionally, the step of determining the overtaking planning path and overtaking vehicle speed that meet the target driving style based on the quadratic programming algorithm includes:

[0018] Plan an initial overtaking path that conforms to the target driving style based on the heuristic algorithm;

[0019] Intercept the local path where the vehicle is located in the initial overtaking path. Take the projection point of the vehicle as the coordinate origin, select points within an appropriate length for smoothing, and use the set of smoothed points as the reference line;

[0020] Determine the cost function value of the reference line based on the preset line trajectory quality cost function;

[0021] Iteratively calculate the overtaking path, and stop when the cost function value is less than the minimum value of the upper boundary of the local convex space or greater than the maximum value of the lower boundary of the local convex space, to obtain the current optimal overtaking path, where the local convex space is a subspace of the complete constraint space composed of multiple local paths;

[0022] Convert the current optimal overtaking path obtained when the iteration stops into Cartesian coordinates to obtain the overtaking planning path;

[0023] An S-T graph is established based on the overtaking planned path, and based on the S-T graph, dynamic programming is used to plan the overtaking speed of the vehicle, where S represents the driving distance of the vehicle along the overtaking planned path, and T represents the length of the driving time.

[0024] Optionally, the expression of the preset line trajectory quality cost function is:

[0025] H(x) = w r f 1 (x,y) + w s f 2 (x,y) + w c f 3 (x,y)

[0026] In the formula, f1(x,y) represents the similarity cost function,

[0027] f2(x,y) represents the smoothness cost function,

[0028] f3(x,y) represents the compactness cost function,

[0029] w r is the weight of the similarity cost, w s is the weight of the smoothness cost; w c is the weight of the compactness cost.

[0030] Optionally, the establishing an S-T graph based on the overtaking planned path and using dynamic programming to plan the overtaking speed of the vehicle based on the S-T graph includes:

[0031] Taking the overtaking planned path as the coordinate axis, establishing a Frenet coordinate system, projecting the vehicle's leading vehicle into the Frenet coordinate system, and establishing an S-T graph, where S represents the driving distance of the vehicle along the overtaking planned path, and T represents the length of the driving time;

[0032] Based on the S-T graph, using dynamic programming to plan the overtaking speed of the vehicle, where the dynamic programming satisfies the leading vehicle distance constraint w o , speed constraint w v and acceleration constraint w a ;

[0033] Among them, taking the distance between the vehicle's head and the leading vehicle's tail as the leading vehicle distance constraint w o ; The expression of the speed constraint w v is:

[0034] w v = f 2 (v - v 0 ) 2

[0035] The acceleration constraint w a is:

[0036] w a = f 3 a 2

[0037] a = (v - v 1 ) × Δt -1

[0038] In the formula, v 0 is the road speed limit, f 2 is the speed cost weight, f 3 is the acceleration cost weight, v is the current vehicle speed, v 1 is the vehicle speed at the planning starting point, and Δt is the time difference between v and v 1 .

[0039] Optionally, the upper - layer speed optimization strategy includes an acceleration strategy, a constant - speed strategy, and a deceleration strategy, where:

[0040] The acceleration process in the acceleration strategy satisfies the following expression:

[0041]

[0042] In the formula, V(t, α) is the final speed of the vehicle acceleration process, V 0 is the initial speed of the vehicle before acceleration; t 0 and t e are the initial and final moments of the acceleration process respectively, and the value range of α is (1, ∞);

[0043] The vehicle speed in the constant - speed strategy is the driving speed of the vehicle at the minimum fuel consumption;

[0044] The deceleration process in the deceleration strategy satisfies the following expression:

[0045]

[0046] In the formula, V is the final speed of the vehicle deceleration process, V 0 is the initial speed of the vehicle deceleration process, γ ≥ 1, t 0 and t e are the initial and final moments of the deceleration process.

[0047] Optionally, the lower - layer depth Q - learning energy management strategy includes state S, action A, reward R, and time T, where:

[0048] The state S includes the remaining charge in the battery and the engine demand power, where the range of the demand power is set to [-80, 80], and the range of the remaining charge in the battery is set to [0, 100];

[0049] The action A is the engine-battery pack power, and the action variable is discretized into 11 actions within the range of [0, 80]:

[0050] A = {0, 10, 20, 30, 40, 50, 60, 70, 80}

[0051] The reward R is set to equivalent fuel consumption and SOC maintenance.

[0052] In addition, to achieve the above object, the present application further provides a vehicle control system, which includes: a memory, a processor, and a vehicle control program for an autonomous driving overtaking scenario that combines path planning and driving style and is stored on the memory and can run on the processor. When the vehicle control program for the autonomous driving overtaking scenario that combines path planning and driving style is executed by the processor, the steps of the vehicle control method for the autonomous driving overtaking scenario that combines path planning and driving style as described above are implemented.

[0053] In addition, to achieve the above object, the present application further provides a computer-readable storage medium, on which a vehicle control program for an autonomous driving overtaking scenario that combines path planning and driving style is stored. When the vehicle control program for the autonomous driving overtaking scenario that combines path planning and driving style is executed by a processor, the steps of the vehicle control method for the autonomous driving overtaking scenario that combines path planning and driving style as described above are implemented.

[0054] The present application has at least the following beneficial effects:

[0055] 1. Through driving simulation experiments and data analysis, identify and cluster the style characteristics of different drivers;

[0056] 2. Based on driver characteristics, construct an overtaking path planning model, and combine vehicle dynamics constraints to optimize the path and speed using dynamic programming and quadratic programming algorithms;

[0057] 3. Use a hierarchical energy management strategy that includes an upper-layer speed optimization strategy and a lower-layer deep Q-learning energy management strategy to ensure optimal energy distribution between the engine and the battery, thereby improving the fuel economy of the vehicle.

[0058] 4. It can reduce energy consumption and significantly improve fuel economy under different driving styles, and greatly enhance the practical application value and popularization potential of autonomous driving technology. Description of the Drawings

[0059] Figure 1 It is a schematic diagram of the architecture of the hardware operating environment of the vehicle control system involved in the embodiments of the present application;

[0060] Figure 2 It is a schematic flowchart of the first embodiment of the vehicle control method for the automatic driving overtaking scenario that combines path planning and driving style in the present application;

[0061] Figure 3 It is a scree plot of the driver style characteristic parameters involved in the embodiments of the present application;

[0062] Figure 4 It is an effect diagram of clustering involved in the embodiments of the present application;

[0063] Figure 5 It is a schematic diagram of the reference line smoothing parameters involved in the embodiments of the present application;

[0064] Figure 6 It is an S-T diagram involved in the embodiments of the present application;

[0065] Figure 7 It is a schematic diagram of the overtaking moment under the S-T diagram involved in the embodiments of the present application;

[0066] Figure 8 It is a schematic diagram of the dynamic programming process under the S-T diagram involved in the embodiments of the present application;

[0067] Figures 9(a)-(c) are respectively schematic diagrams of a vehicle overtaking a single vehicle in front, a vehicle overtaking a traffic flow in front, and a vehicle overtaking in a curve section involved in the embodiments of the present application;

[0068] Figure 10 It is a schematic diagram of the relationship between speed-time and acceleration-time corresponding to different acceleration characteristic parameters involved in the embodiments of the present application;

[0069] Figure 11 It is a schematic diagram of the speed curves of an actual vehicle corresponding to different deceleration characteristic parameters involved in the embodiments of the present application;

[0070] Figure 12 It is a schematic diagram of the DQN principle involved in the embodiments of the present application;

[0071] Figures 13(a)-(c) are respectively schematic diagrams of the SOC change curve and fuel consumption of overtaking by a conservative driver style before optimization, the SOC change curve and fuel consumption of overtaking by a normal driver style before optimization, and the SOC change curve and fuel consumption of overtaking by an aggressive driver style before optimization involved in the embodiments of the present application;

[0072] Figures 14(a)-(c) are respectively the schematic diagrams of the SOC change curves and fuel consumption of the optimized conservative driver style overtaking, the optimized normal driver style overtaking, and the optimized aggressive driver style overtaking involved in the embodiments of the present application;

[0073] Figure 15 It is the schematic diagram of the SOC change curve and fuel consumption before distinguishing the driver style involved in the embodiments of the present application.

[0074] The realization, functional features, and advantages of the purpose of the present application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed Embodiments

[0075] To better understand the above technical solutions, the exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0076] As an implementation solution, Figure 1 It is the schematic diagram of the architecture of the hardware operating environment of the vehicle control system involved in the embodiment solution of the present application.

[0077] As Figure 1 shown, the vehicle control system may include: a processor 1001, such as a CPU, a memory 1005, a user interface 1003, a network interface 1004, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0078] Those skilled in the art can understand that Figure 1 the vehicle control system architecture shown in does not constitute a limitation on the vehicle control system and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0079] AsFigure 1 As shown in Figure 1 , the memory 1005 as a storage medium may include an operating system, a network communication module, a user interface module, and a vehicle control program for an autonomous overtaking scenario that combines path planning and driving style. Among them, the operating system is a program that manages and controls the hardware and software resources of the vehicle control system, and combines the operation of the vehicle control program for the autonomous overtaking scenario that combines path planning and driving style and other software or programs.

[0080] In Figure 1 In the vehicle control system shown in Figure 1 , the user interface 1003 is mainly used to connect to the terminal and communicate with the terminal for data; the network interface 1004 is mainly used for the background server and communicates with the background server for data; the processor 1001 can be used to call the vehicle control program for the autonomous overtaking scenario that combines path planning and driving style stored in the memory 1005.

[0081] In this embodiment, the vehicle control system includes: a memory 1005, a processor 1001, and a vehicle control program for an autonomous overtaking scenario that combines path planning and driving style and is stored on the memory and can run on the processor, where:

[0082] When the processor 1001 calls the vehicle control program for the autonomous overtaking scenario that combines path planning and driving style stored in the memory 1005, the following operations are performed:

[0083] Obtain the historical driving characteristic parameters of the vehicle, and determine the target driving style type according to the historical driving characteristic parameters. The historical driving characteristic parameters include the mean lateral acceleration, the mean lateral speed, the steering wheel angle, the standard deviation of the lateral speed, the standard deviation of the lateral acceleration, and the average acceleration. The target driving styles include an aggressive driving style, a normal driving style, and a conservative driving style;

[0084] Based on the quadratic programming algorithm, determine the overtaking planning path and the overtaking vehicle speed that meet the target driving style;

[0085] Based on the overtaking planning path, the overtaking vehicle speed, and the target driving style, formulate an upper-layer speed optimization strategy and a lower-layer deep Q-learning energy management strategy;

[0086] Control the vehicle speed constraint during overtaking to meet the upper-layer speed optimization strategy, and control the power distribution between the vehicle's battery and engine to meet the lower-layer deep Q-learning energy management strategy.

[0087] When the processor 1001 calls the vehicle control program for the autonomous overtaking scenario that combines path planning and driving style stored in the memory 1005, the following operations are performed:

[0088] Dimensionality reduction is performed on the historical driving characteristic parameters based on the principal component analysis method to form a preprocessed data set;

[0089] Calculate the Euclidean distance between each point in the preprocessed data set and the center point;

[0090] Based on the Euclidean distance, perform K-means clustering to calculate the center point value of the lateral speed, the center point value of the lateral acceleration, and the center point value of the steering wheel mean;

[0091] Determine the driving style type corresponding to the interval where the center point value of the lateral speed, the center point value of the lateral acceleration, and the center point value of the steering wheel mean are located as the target driving style type.

[0092] When the processor 1001 calls the vehicle control program for the automatic driving overtaking scenario that combines path planning and driving style stored in the memory 1005, the following operations are performed:

[0093] Based on the heuristic algorithm, plan an initial overtaking path that conforms to the target driving style;

[0094] Intercept the local path where the vehicle is located in the initial overtaking path. Taking the projection point of the vehicle as the coordinate origin, select points within an appropriate length for smoothing, and the set of smoothed points is used as the reference line;

[0095] Determine the cost function value of the reference line based on the preset line trajectory quality cost function;

[0096] Iteratively calculate the overtaking path, and stop when the cost function value is less than the minimum value of the upper boundary of the local convex space or greater than the maximum value of the lower boundary of the local convex space, to obtain the current optimal overtaking path, where the local convex space is a subspace of the complete constraint space composed of multiple local paths;

[0097] Convert the current optimal overtaking path obtained when the iteration stops into Cartesian coordinates to obtain the overtaking planned path;

[0098] Based on the overtaking planned path, establish an S-T graph, and based on the S-T graph, use dynamic programming to plan the overtaking vehicle speed of the vehicle, where S represents the driving distance of the vehicle along the overtaking planned path, and T represents the length of the driving time.

[0099] Based on the hardware architecture of the vehicle control system based on the above automatic driving technology, an embodiment of the vehicle control method for the automatic driving overtaking scenario that combines path planning and driving style in the present application is proposed.

[0100] The first embodiment

[0101] Refer to Figure 2, in the first embodiment, the vehicle control method for an autonomous driving overtaking scenario that combines path planning and driving style includes the following steps:

[0102] Step S10, obtain the historical driving characteristic parameters of the vehicle, and determine the target driving style type according to the historical driving characteristic parameters. The historical driving characteristic parameters include the mean lateral acceleration, the mean lateral speed, the steering wheel angle, the standard deviation of the lateral speed, the standard deviation of the lateral acceleration, and the average acceleration. The target driving styles include an aggressive driving style, a normal driving style, and a conservative driving style;

[0103] In this embodiment, the mean lateral acceleration, the mean lateral speed, the steering wheel angle, the standard deviation of the lateral speed, the standard deviation of the lateral acceleration, and the average jerk are used as characteristic parameters to identify the driver's style. Among them, the average jerk can be interpreted as the jerk degree, and through the jerk degree, the influence of the shift shock on the driver's feeling degree can be truly reflected.

[0104] Optionally, the step of determining the target driving style type according to the historical driving characteristic parameters includes: performing dimensionality reduction on the historical driving characteristic parameters based on the principal component analysis method to form a preprocessing data set; calculating the Euclidean distance between each point in the preprocessing data set and the center point; performing K-means clustering based on the Euclidean distance, and calculating the center point value of the lateral speed, the center point value of the lateral acceleration, and the center point value of the mean steering wheel; determining the driving style type corresponding to the interval where the center point value of the lateral speed, the center point value of the lateral acceleration, and the center point value of the mean steering wheel are located as the target driving style type.

[0105] In this embodiment, in order to accurately classify the driver's style in the overtaking environment, drivers are recruited for driving simulation tests. After extracting the vehicle driving data, the overtaking process is decomposed. Combining the research, typical characteristic parameters of the overtaking process are selected. The principal component analysis method is used to perform dimensionality reduction on the parameters. Combining the contribution degree and the scree plot, the main characteristic parameters are selected. Further, the K-means clustering method is used to identify the driver's style with the main characteristic parameters as indicators, and it is divided into an aggressive type, a normal type, and a conservative type. And the silhouette coefficient, CH, and DBI are used as three indicators to evaluate the clustering effect of the driver's style.

[0106] Identifying the driver's style based on the selected characteristic parameters ensures the accuracy of driver's style identification. However, it has a large computational load, and there is a certain correlation between the parameters, which will have a certain impact on driver's style identification. In statistics, Principal Component Analysis (PCA) is universal for data dimensionality reduction. In this embodiment, there are many characteristic parameters for driver's style classification, and PCA can be used to simplify and reduce the dimensions of the parameters. Dimensionality reduction is a method for preprocessing high-dimensional feature data. While retaining the features of high dimensions, PCA can also remove noise and reduce the computational overhead of the algorithm, thereby achieving the purpose of improving the data processing speed and saving a large amount of time costs.

[0107] It should be noted that the core idea of PCA is to map n-dimensional features to k dimensions. By calculating the covariance matrix of the data, the eigenvalue eigenvectors of the covariance matrix are obtained, and a matrix composed of the eigenvectors corresponding to the k largest eigenvalues is selected. The covariance can be expressed as:

[0108]

[0109] In the formula, is the mean. When the covariance is positive, it indicates that X and Y are positively correlated; on the contrary, it is negatively correlated. When the covariance is 0, it indicates that X and Y are independent. When the sample is n-dimensional data, their covariance is actually a covariance matrix, that is:

[0110]

[0111] Since the selected parameters have different dimensions, the differences between the parameters are not conducive to PCA. Therefore, it is necessary to normalize them, convert the parameter eigenvalues to the same dimension, and map the data to the interval [0,1]. The driver's style characteristic parameters can be represented by the matrix X as:

[0112]

[0113] Standardizing the matrix X can obtain the standardized mean Z. Using the standardized matrix Z to calculate the covariance matrix R, then:

[0114]

[0115] Then the eigenvalue λ of the covariance matrix R can be expressed as:

[0116] λ = (λ 1 , λ 2 , λ 3 … λ n )

[0117] The cumulative contribution rate can be calculated using the first m eigenvalues. When the cumulative contribution rate exceeds 90%, it can be considered that the selected principal components can represent all components. The variance explanation rate and cumulative explanation rate of each component are calculated using SPSS software. SPSS software is a combined software package and the world's earliest statistical software with a graphical menu-driven interface. It can input, organize, and analyze data. As the earliest statistical software with a graphical menu-driven interface, its operation interface is quite user-friendly, the output results are relatively intuitive, and it has low requirements for the computer hardware system and a relatively loose requirement for the operating environment, fully meeting the data processing needs of non-statistical professionals. Therefore, it has played an important role in many fields of social science and natural science.

[0118] Perform PCA calculations using SPSS software. The variance contribution rate and cumulative contribution rate of each parameter obtained are shown in Table 2.1. In the table, principal component 1 represents the mean lateral speed, principal component 2 represents the mean lateral acceleration, principal component 3 represents the mean steering wheel angle, principal component 4 represents the standard deviation of the lateral speed, principal component 5 represents the standard deviation of the lateral acceleration, and principal component 6 represents the mean jerk. The scree plot is as Figure 3 shown.

[0119] Table 1. Variance explanation rate and cumulative explanation rate of each component

[0120]

[0121] As can be seen from Table 1, the variance explanation rate of principal component 1 is 68.056%, and the cumulative explanation rate is 68.056%; the variance explanation rate of principal component 2 is 19.88%, and the cumulative explanation rate is 87.936%; the variance explanation rate of principal component 3 is 9.117%. The cumulative explanation rate of the first three components reaches 95.201%, exceeding 95%. This shows that the mean lateral speed, mean lateral acceleration, and mean steering wheel can well reflect the original driving behavior information. Therefore, in this embodiment, the mean lateral speed, mean lateral acceleration, and mean steering wheel angle are selected as the characteristic parameters for driver style recognition for further analysis.

[0122] Furthermore, in this embodiment, the Euclidean distance is used as the distance from the points in the dataset to the center point. The formula for calculating the Euclidean distance is as follows:

[0123]

[0124] In the formula, d is the distance from each data to the center point; (x, y) is the position of each data.

[0125] In all datasets, K data points are taken as the clustering centers for each cluster. In this embodiment, 3 clustering centers are set, that is, K = 3, denoted as Means1, Means2, and Means3. After setting the number of clustering centers, iteration is performed, and its iteration formula is as shown in Equation 2.7:

[0126]

[0127] In the formula, c (j) is the clustering center closest to each data; x (i) -Mean(j) is the Euclidean distance; r (ij) If the value is attributed to c (j) it is 1, otherwise it is 0.

[0128] Exemplarily, K-means clustering is performed using Python software, and the clustering effect is as shown in Figure 4 It can be seen that SA represents the mean lateral speed, AA represents the mean lateral acceleration, and SWA represents the mean steering wheel angle. According to Figure 4 it can be known that during the overtaking process, the number of normal and conservative drivers is basically the same, each accounting for 38.9% of the total number of people. The lateral acceleration, lateral speed, and steering wheel angle of most drivers are relatively small, indicating that most people first pay attention to overtaking safety during the entire overtaking process. The proportion of aggressive drivers is relatively small, about 22.2%, which means that although the number of aggressive drivers is small during the overtaking process, they are more sensitive to speed and are more likely to have sudden acceleration and deceleration behaviors. The coordinates of the clustering center points are shown in Table 2.3. It is found that the indicators of conservative drivers are all relatively small, which is in line with the understanding of the conservative style. The values of normal drivers are all in the middle, which provides a scientific basis for the normal driver style as a transitional driver style. The parameters of aggressive drivers are all relatively high, indicating that aggressive drivers pay more attention to overtaking efficiency during the entire overtaking process.

[0129] Furthermore, the driving style type corresponding to the interval where the central point value of the lateral speed, the central point value of the lateral acceleration, and the central point value of the steering wheel mean are located is determined as the target driving style type. Optionally, the corresponding interval can refer to Table 2 below:

[0130] Table 2 Coordinate Table of Clustering Center Points

[0131]

[0132] In addition, it should also be noted that since the K-means clustering method belongs to unsupervised learning and does not label the clustering parameters, the clustering effect cannot be quantified, but the clustering effect can be evaluated according to the density and dispersion within the clusters.

[0133] The density is measured by calculating the average distance between data points within each category. When the data points within a category are closer, the average distance is smaller, indicating a better clustering effect. The dispersion is measured by calculating the distance between the centroids of different clusters. If the distance between the centroids of each cluster is larger, it indicates a better clustering effect. The Davies-Bouldin Index (DBI) is used to evaluate the number of clusters K. The core idea of DBI is to calculate the similarity between each cluster. The smaller the value of DBI, the lower the dispersion and the better the clustering effect. The silhouette coefficient is used to evaluate the quality of the clustering effect. The silhouette coefficient is an index that describes the clarity of the silhouette of each cluster after clustering. It is measured by calculating the distance between each data point and other data points within its cluster and the distance to the centroid of the nearest other cluster, including two factors: cohesion and separation. The expression formula of the silhouette coefficient is shown as follows:

[0134] S (i) =(b i -a i )(max(a i ,b i )) -1

[0135] In the formula, a i and b i represent cohesion and separation, where the calculation formula of a i is as follows:

[0136]

[0137] In the formula, D(i,j) represents the distance between sample point i and other point j. The smaller a(i) is, the closer the clustering is. The calculation method of b i is similar to that of a i . The difference is that a i calculates the distance between points within the same cluster, while b i calculates the distance to points in other clusters. Therefore, it can be rewritten as:

[0138]

[0139] The value range of the silhouette coefficient S is [-1, 1]. When a(i) < b(i), the within-class distance is less than the between-class distance, the clustering result is more compact, the silhouette coefficient S approaches 1, and the sample has a high similarity with other samples in its own cluster; when a(i) = b(i), the silhouette coefficient S is 0, indicating that the two samples belong to the same cluster; when a(i) > b(i), it means that the within-class distance is greater than the between-class distance, the silhouette coefficient S approaches -1, the clustering result is relatively loose, and the clustering effect is worse. According to the clustering result, calculate using the above silhouette coefficient calculation method, and take the average value of the silhouette coefficients of all samples as the evaluation index. The evaluation index table is shown in Table 3 below:

[0140] Table 3 Evaluation Index Table

[0141]

[0142] As can be seen from Table 3, the average value of the silhouette coefficients of all samples S > 0, indicating that the clustering effect is good. CH (Calinski-Harbasz Score) represents the ratio of the compactness to the separation degree. The larger the CH, the greater the distance between clusters and the higher the within-cluster compactness, and the better the clustering effect. It can be seen from the above table that the CH value in the clustering effect evaluation index of this embodiment is 19.082, indicating that the clustering effect is relatively ideal.

[0143] Step S20, based on the quadratic programming algorithm, determine the overtaking planning path and overtaking speed that meet the target driving style;

[0144] In this embodiment, based on the driver characteristics, an overtaking path planning model is constructed, and combined with the vehicle dynamics constraints, the dynamic programming and quadratic programming algorithms are used to optimize the path and speed.

[0145] Quadratic optimization is an extremely important optimization algorithm in optimal problems. In this embodiment, the quadratic optimization algorithm is introduced to solve the vehicle overtaking trajectory. During the entire overtaking process of the vehicle, the obstacle avoidance space for overtaking the vehicle in front becomes a typical non-convex problem. To solve this problem, this embodiment first discretizes the trajectory, and uses a heuristic algorithm to search for a rough solution of the overtaking path. Based on the rough solution, a convex space is opened up, and the optimal solution is optimized in the convex space. That is, various constraints of the path planning model are used to form a high-dimensional non-convex space, and a local convex space is constructed in this space. Since the local convex space is a subspace of the complete constraint space, the solution obtained under the local convex space can meet the constraints. Since there are multiple paths for the rough solution, it is transformed into a shortest path problem. At the same time, the DP algorithm is used to find the rough solution in the discrete space, the five-degree polynomial is used to connect the discrete points of the DP algorithm, and the quadratic programming is used to solve the final solution.

[0146] After calculating the optimal overtaking trajectory using quadratic programming, speed planning is performed on the vehicle, serving as the research basis for subsequent fuel economy. Speed planning determines the entire process of the vehicle's overtaking trajectory and is also affected by path planning. Therefore, it also needs to satisfy collision constraints and speed limits on urban roads.

[0147] Step S30: Based on the overtaking planned path, the overtaking vehicle speed, and the target driving style, formulate an upper-layer speed optimization strategy and a lower-layer deep Q-learning energy management strategy.

[0148] In this embodiment, a hierarchical optimization control strategy is proposed. The upper layer combines the driver's style to design speed optimization strategies for three working conditions: acceleration, deceleration, and constant speed, optimizes the vehicle speed based on different driver styles, and explores the relationship between the speed optimization curve and vehicle fuel economy. At the same time, two dimensions of driving safety and comfort are further considered. The lower layer constructs a deep Q-learning model. After briefly introducing the principle of the deep Q-learning model, relevant parameters of the deep Q-learning are adjusted. Combining the speed optimization method, vehicle speed information of different driver styles, and the equivalent fuel consumption minimum strategy, an energy management strategy based on different driver styles is established. The effectiveness of the proposed speed optimization method and energy management strategy is verified by result comparison, providing a scientific basis for the fuel economy of the overtaking process during the automatic driving of hybrid vehicles.

[0149] Step S40: Control the vehicle speed constraint during overtaking to satisfy the upper-layer speed optimization strategy, and control the power distribution between the vehicle's battery and engine to satisfy the lower-layer deep Q-learning energy management strategy.

[0150] Finally, control the vehicle speed constraint during overtaking to satisfy the upper-layer speed optimization strategy, and control the power distribution between the vehicle's battery and engine to satisfy the lower-layer deep Q-learning energy management strategy.

[0151] Exemplarily, in some specific embodiments, vehicle parameters are set, mainly including parameters such as the power of the engine and the motor. The specific parameter settings are shown in Table 4, where the peak power of the engine is set to 80 kw, the peak power of the motor is set to 60 kw, and the vehicle mass is set to 2100 Kg.

[0152] Table 4. Main vehicle parameters

[0153]

[0154] Second Embodiment

[0155] Based on the first embodiment, in this embodiment, step S20 includes:

[0156] Step S21, plan an initial overtaking path that conforms to the target driving style based on a heuristic algorithm;

[0157] Step S22, intercept the local path where the vehicle is located in the initial overtaking path, take the projection point of the vehicle as the coordinate origin, select points within an appropriate length for smoothing, and use the set of smoothed points as the reference line;

[0158] Step S23, determine the cost function value of the reference line based on a preset line trajectory quality cost function;

[0159] Step S24, iteratively calculate the overtaking path, and stop when the cost function value is less than the minimum value of the upper boundary of the local convex space or greater than the maximum value of the lower boundary of the local convex space, to obtain the current optimal overtaking path, where the local convex space is a subspace of the complete constraint space composed of multiple local paths;

[0160] Step S25, convert the current optimal overtaking path obtained when the iteration stops into Cartesian coordinates to obtain the overtaking planned path;

[0161] Step S26, establish an S-T graph based on the overtaking planned path, and based on the S-T graph, use dynamic programming to plan the overtaking vehicle speed, where S represents the driving distance of the vehicle along the overtaking planned path, and T represents the length of the driving time.

[0162] In this embodiment, to search for the optimal path in the convex space, it is first necessary to generate a reference line because the global path is too long, which is not conducive to coordinate conversion and may lead to non-unique projection points of the vehicle in front. The reference line can solve the problem of the navigation path being too long and not smooth, that is, intercept the local path where the vehicle is located in the global path, take the projection point of the vehicle as the coordinate origin, select points within an appropriate length for smoothing, and use the set of smoothed points as the reference line.

[0163] The smoothness of the reference line is related to the road alignment. The closer the road alignment is to a straight line, the smoother the reference line. However, being close to a straight line may lead to a large difference from the original path. As Figure 5 shown, in the figure, Pi (i = 1, 2, 3) represents the points in the original global path, and Ri (i = 1, 2, 3) represents the smoothed reference line.

[0164] From Figure 5 it is found that the difference between the smoothed line segment P1P3 and the original road segment is large. It is found that the smaller P1R1 + P2R2 + P3R3 is, the closer the smoothed line segment is to the original path. In order to better couple the smoothed path with the reference line, therefore, the expression of the preset line trajectory quality cost function is obtained as:

[0165] H(x) = w r f1 (x, y) + w s f 2 (x, y) + w c f 3 (x, y)

[0166] In the formula, f1(x, y) represents the similarity cost function,

[0167] f2(x, y) represents the smoothing cost function,

[0168] f3(x, y) represents the compactness cost function,

[0169] w r is the weight of the similarity cost, w s is the weight of the smoothing cost; w c is the weight of the compactness cost.

[0170] The smaller the cost function, the smoother and more similar the reference line is. It can be known that the function is a typical quadratic programming problem. It is found that the cost function belongs to the quadratic programming problem, and the quadratic programming can be used to solve it. Quadratic programming is a classic form of convex optimization, and its general expression is shown as follows:

[0171]

[0172] s.t. Ax ≤ b

[0173] lb ≤ x ≤ ub

[0174] In the formula, Q is the Hessian matrix. When the Hessian matrix is a positive semi - definite matrix, the above function is a convex quadratic programming; when the Hessian matrix is a positive definite matrix, the function is a strictly convex quadratic programming and has a unique global minimum. At the same time, the quadratic programming satisfies two constraints, including the collision constraint and the planning starting point constraint, and Equation 4.22 is transformed into a quadratic programming problem for solution.

[0175] Expanding the similarity cost f1(x, y), it is found that the function is not affected by the sum of squares of the reference line coordinates. The similarity cost function can be calculated from the matrix composed of smoothing coordinates and the identity matrix, that is:

[0176] f 1 (x, y) = (x 1 , y 1 ,..., x n , y n )A 1 (x 1 , y 1 ,..., x n , y n ) T

[0177] -2(x r1 ,y r1 ,...,x rn ,y rn )(x 1 ,y 1 ,...,x n ,y n ) T

[0178] Let \((x 1 ,y 1 ,x 2 ,y 2 ,...,x n ,y n ) be denoted as \(x^T\), \(A_1\) be the identity matrix, and the product of the reference line coordinates and the constant term be denoted as the function \(M(x)\). Then the similarity cost function \(h_1(x,y)\) can be transformed into:

[0179] h 1 (x,y)=w r x T A 1 T A 1 x - Mx

[0180] Transform the smoothing cost \(f_2(x,y)\) into matrix form. Then, for \(n\) points, there are \(n - 2\) terms, and the matrix will be continuously repeated diagonally in the form of a \(6\times2\) sub - matrix. Finally, the matrix is \(2n\) rows and \(2n - 4\) columns, as shown in Equation 4.26:

[0181]

[0182] Denote the matrix as \(A_2^T\). Then the smoothing cost is:

[0183] h 2 (x,y)=w s x T A 2 T A 2 x

[0184] Similarly, the compactness cost \(h_3(x,y)\) can be transformed into:

[0185] h 3 (x,y)=w c x T A 3 T A 3 x

[0186] Transform the penalty function into the quadratic programming form. Combining the above - mentioned equations, it can be seen that the final quadratic programming form is:

[0187] Q = 2(w r A 1 T A 1 + w s A 2 T A 2 + w c A 3 T A 3 )

[0188] c T = w r M

[0189] The final path planning result is obtained through multiple iterations. Each iteration process is a convex quadratic programming. The iteration can be terminated when it satisfies within the rough solution of dynamic programming, and the four vertices of the vehicle are less than the minimum value of the upper boundary of the convex space and greater than the maximum value of the lower boundary of the convex space. Finally, it is transformed into Cartesian coordinates to obtain the final path planning result.

[0190] On the other hand, in urban roads, when the vehicle in front of the ego vehicle cannot reach the desired speed, there are usually two decisions. The first is to accelerate to overtake the vehicle or the traffic flow in front, and the second is to decelerate and follow the vehicle or the traffic flow in front. Since this embodiment mainly studies the overtaking scenario on urban roads, this embodiment focuses on studying the first decision and does not consider the speed planning of the vehicle in the following behavior scenario. Due to the complex game problem between the ego vehicle and the vehicle in front during the entire overtaking process, in order to simplify the problem in overtaking speed planning, this embodiment assumes that the vehicle in front travels at a constant speed and a constant acceleration. Taking the planned overtaking trajectory as the coordinate axis, after establishing the Frenet coordinate system, the vehicle in front is projected into the Frenet coordinate system, and an S-T diagram is established accordingly, as Figure 6 shown. The S-T diagram mainly considers the geometric shapes and spatio-temporal relationships of the ego vehicle and the vehicle in front, projects the vehicle in front onto the path planned by the ego vehicle, that is, when the vehicle in front will have a spatio-temporal collision with the ego vehicle, and when the ego vehicle can start and complete the overtaking process. Referring to Figure 7 the overtaking moment schematic diagram shown in, where S represents the distance traveled by the vehicle along the planned path, T represents the length of the traveled time, s1 represents that the rear of the ego vehicle just overtakes the vehicle in front at time t1, and s2 represents that the ego vehicle completes the overtaking behavior and returns to its own lane to continue driving at time t2.

[0191] The complete trajectory space contains four dimensions: x, y, θ, and t. It is difficult to directly perform speed planning. By using S-T, the planning dimension is reduced to two dimensions, and the speed planning problem is transformed into a problem similar to path planning in the Frenet coordinate system for solution. Appropriate points are selected from the planned path length as the coordinate axes of the Frenet coordinate system, and the vehicle ahead is projected onto this coordinate system.

[0192] Dynamic programming and quadratic programming are used to plan the speed. First, the starting point is planned. The driving information of the current vehicle is determined using the information of the vehicle at the previous moment before the planned starting point. Since the path planning determines the coordinates, yaw angle, speed, and acceleration of the vehicle at the previous moment, the S of the vehicle at the planned starting point is 0. Here, S is the offset distance of the vehicle, is the magnitude of the speed V, and is the magnitude of the acceleration. According to research, the vehicle can complete the overtaking process within 10 s, and the vehicle can complete the process of overtaking the vehicle ahead and returning to the original lane within 3 s. Then, the process of dynamic programming within 10 s is as Figure 8 shown:

[0193] Different from the DP in path planning, this DP does not use a fifth-degree polynomial connection but directly uses a straight line connection because the connection between different layers of the DP in path planning represents the spatial driving trajectory of the vehicle, and the discrete point spacing is large, while the DP in speed planning represents the speed on the trajectory, and the discrete point spacing is small. Using DP to plan the speed needs to satisfy three constraints, namely the leading vehicle distance constraint wo, the speed constraint wv, and the acceleration constraint wa.

[0194] Among them, the distance between the front of the vehicle and the rear of the leading vehicle is used as the leading vehicle distance constraint w o ; the expression of the speed constraint w v is:

[0195] w v = f 2 (v - v 0 ) 2

[0196] The acceleration constraint w a is:

[0197] w a = f 3 a 2

[0198] a = (v - v 1 ) × Δt -1

[0199] In the formula, v 0 is the road speed limit, f 2 is the speed cost weight, f 3is the acceleration cost weight, v is the current vehicle speed, v 1 is the vehicle speed at the planning starting point, and Δt is the time difference between v and v 1

[0200] In each layer of the DP planning, starting from the minimum values of the distance, speed, and acceleration from the first layer to the planning starting point, iterative calculations are performed downward until the vehicle returns to the center line of the original lane. The DP planning completes the iteration, and the minimum cost is used as the output value of the speed planning and the input value of the quadratic programming. The cost function is solved through quadratic programming.

[0201] In addition, in this embodiment, to verify the feasibility of the path planning algorithm based on the convex optimization algorithm, in a specific implementation, joint simulation is performed using PreScan / CarSim / SimuLink, and the path planning results are visualized using Matlab. The processor is inter(R)core(TM)2.90GHz, and the operating system is 64-bit.

[0202] In the simulation environment, the vehicle width is set to 3.5m, and there are two-way two lanes. In the simulation result diagram, the blue rectangular frame represents the rectangular contour of the ego vehicle, the black rectangular frame represents the rectangular contour of the vehicle ahead sensed by the ego vehicle's sensor, and the red line represents the overtaking trajectory planning result. The overtaking environment on urban roads is simulated and tested. According to the common overtaking scenarios on urban roads, it is divided into 3 types, specifically including overtaking a single vehicle ahead, overtaking a traffic flow ahead, and overtaking on a curve section, as shown in Figures 9(a)-(c).

[0203] It can be seen that when the ego vehicle starts to drive on a straight section and there is a vehicle or traffic flow ahead in the current lane, the speed of the vehicle ahead cannot reach the desired speed of the ego vehicle, and at the same time, the conventional lane-changing is difficult to meet the obstacle avoidance requirements. To avoid collision with the vehicle ahead, it is necessary to change lanes and overtake, switch to the left lane to overtake, and return to the original lane to continue driving after overtaking. It can be found that when the vehicle encounters a vehicle or traffic flow ahead during driving, the path can be planned in advance through the algorithm built in this embodiment, and the overtaking can be completed safely and the vehicle can return to the original lane to continue driving.

[0204] Third Embodiment

[0205] In this embodiment, based on the first embodiment, the upper-layer speed optimization strategies include an acceleration strategy, a constant-speed strategy, and a deceleration strategy.

[0206] Among them, the acceleration process in the acceleration strategy satisfies the following expression:

[0207]

[0208] In the formula, V(t, α) is the final speed of the vehicle during the acceleration process, and V 0 is the initial speed of the vehicle before acceleration; t​0 and t e are the initial and final moments of the acceleration process respectively, and the value range of α is (1, ∞).

[0209] Exemplarily, taking the impulsive driver style as an example, a driver with an impulsive driver style is prone to quickly stepping on the pedal and the accelerator when driving. To design an energy-saving strategy suitable for impulsive drivers, an "acceleration - deceleration" strategy is proposed. Combining with the maximum speed limit on urban roads, the speed is limited within 60 Km·h-1. That is, under the operation of an impulsive driver, the vehicle accelerates to the maximum road speed of 60 Km·h-1 and then decelerates.

[0210] Take α = 1.2, 1.4, 1.6, 1.8, 2.0, 2.5, 3.0, 3.5. Refer to Figure 10 the schematic diagrams of the speed-time and acceleration-time relationships corresponding to different acceleration characteristic parameters shown, it can be seen that the larger α is, the more concave the speed curve trend is. That is, as α increases, the acceleration change in the initial stage of the acceleration process gradually slows down, the change rate of the acceleration gradually increases, and the acceleration at the end of the acceleration process is also larger.

[0211] Using the typical NEDC working condition, simulate the vehicle acceleration process with a speed from 0 to 40 Km·h-1 within 30 s, and obtain the simulated values of the driving mileage and energy consumption per 100 kilometers under different α values, as shown in Table 5:[[]]END]]

[0212] Table 5. Energy consumption corresponding to different α values

[0213]

[0214] As α increases, the driving mileage of the vehicle decreases, and the energy consumption per 100 kilometers has a positive correlation with α, that is, the energy consumption per 100 kilometers increases as α increases. This is because the larger α is in the early stage of the acceleration process, the smaller the pressure generated by fuel combustion, resulting in a smaller output power of the engine. Therefore, the vehicle speed and acceleration are smaller in the early stage. As time t increases, the engine power gradually becomes larger, and the acceleration gradually increases, but the change in this process is relatively slow. Therefore, as the vehicle travels at a lower speed for a longer time, its driving mileage gradually decreases, and the energy consumption per 100 kilometers gradually increases.

[0215] For the constant speed strategy, the vehicle maintaining a constant speed has a better energy-saving effect compared to acceleration and deceleration. This speed is called the economic speed. Based on this, the economic speed is defined as the driving speed of the vehicle at the minimum fuel consumption. Taking the urban road signal intersection as the research object, when driving in the urban area, the 3rd gear speed is often used, and the average value of 36 Km·h-1 is taken as the economic speed, and this speed is used as the driving speed of the constant speed strategy.

[0216] For the deceleration strategy, when the host vehicle is traveling at medium or high speed, if the leading vehicle cannot reach the desired speed of the host vehicle, in order to comply with traffic rules and ensure driving safety, the host vehicle needs to take deceleration measures in advance. The forces acting on the vehicle during deceleration are shown in the following equation:

[0217]

[0218] In the formula: F t is the traction force; F R is the tire resistance; F L is the air resistance; F St is the gradient resistance; F a is the acceleration resistance. No traction force is generated during the deceleration process. Assuming the vehicle is traveling on a flat road under the same ground conditions, so F t = 0, F St = 0, and the vehicle speed is exponentially related to F L . The lower the vehicle speed, the smaller F L . Therefore, the time in the early stage of the deceleration process is short while the distance is long, and the time in the later stage of deceleration is long while the distance is short. Combining trigonometric functions and exponents, the deceleration characteristic parameter γ is designed to represent the deceleration process, as shown in the following equation:

[0219]

[0220] In the formula, V is the final speed during the vehicle deceleration process, V 0 is the initial speed during the vehicle deceleration process, γ ≥ 1, t 0 and t e are the initial and final moments of the deceleration process.

[0221] Furthermore, in order to verify the effectiveness of the speed optimization method during the deceleration process, an experimental process was designed and 50 drivers were recruited for the experiment. A 14-second deceleration process was randomly selected from the driving data collected from the 50 drivers and averaged. The ages of the drivers were all between 20 and 50 years old, and their actual driving experience was between 1 and 20 years. Points were taken at intervals of 0.5. As shown in the schematic diagram of the actual vehicle and the corresponding speed curves Figure 11 , it was found that the actual vehicle deceleration curve was between 1.5 and 2.0.

[0222] Fourth Embodiment

[0223] In this embodiment, based on the first embodiment, the lower-layer depth Q-learning, also known as the DQN algorithm (Deep Q Network), mainly uses the good generalization ability of the neural network to transform the update of its own table (Q-table) into a function fitting problem. By using the neural network to replace the Q-table to generate Q values, the optimal action selection for the state s can be obtained, effectively solving the action decision problem in the multi-dimensional state space. The DQN algorithm updates the network with the target value yt, as shown in the following equation:

[0224] y t = r + γmax a′ Q(s t+1 , a′; θ - )

[0225] Where y t is used as the approximate true value; max a′ Q(s t+1 , a′; θ - ) represents obtaining the Q-value by selecting the action a′ that maximizes the Q-value in the state s t+1 ; r represents the reward value; γ represents the discount factor; the difference between the predicted value and the true value is used as the loss function, and the expression is as shown in the following formula:

[0226] Loss(θ) = E[(y t - Q(s t , a t ; θ)) 2

[0227] The DQN algorithm adopts two key techniques, namely experience replay and fixed Q-target network. Experience replay is one of the core ideas of the DQN algorithm. Its basic principle is to store the agent's experiences in a replay memory bank and then randomly sample from it to update the model using these experiences. The advantage of doing this is to avoid the correlation between samples and improve the stability and convergence speed of the model. Fixed Q-target network: The DQN algorithm uses two neural networks. One is the Q-value network (evaluate network), which is used to select actions and update the model; the other is the target network (target network). In Equation 5.4, θ - represents the parameters of the target network and is used to calculate the target Q-value. The two network structures are exactly the same. The difference is that the Q-value network mainly calculates the Q-value of the policy and the update iteration of the Q-value, while the target network is mainly used to calculate the Q-value of the next state, which can reduce the fluctuation of the target and improve the stability of the model.

[0228] Furthermore, the core of the DQN algorithm lies in approximating the Q-value predicted in the current state with the Q-value based on past experiences. The Q-value network of the DQN function is used to calculate the Q-value of policy selection and the iterative update of the Q-value. The DQN algorithm mainly consists of four parameters, including state S, action A, reward R, and time T. The basic principle of the DQN algorithm is to construct a Q-table of state S - action A, and use a neural network to map the state S to the Q-values of all actions that can be executed from this state. That is, when the current state S is input, the network will output the Q-values corresponding to all currently executable actions respectively. The basic principle is as Figure 12 shown in the DQN principle diagram.

[0229] ​Since this embodiment mainly applies the DQN algorithm to allocate power between the battery and the engine, the settings of the state variables, action variables, and reward function of DQN affect the optimization effect of energy management based on deep reinforcement learning. The specific settings are as follows:

[0230] The state variables are set as the required power and SOC. The range of the required power is set as [-80, 80], and the range of SOC is set as [0, 100]. The expression of the state variable S is:

[0231] S = {SOC, P re}

[0232] The action variable is the engine-battery pack power. The action variable is discretized into 11 actions within the range of [0, 80]:

[0233] A = {0, 10, 20, 30, 40, 50, 60, 70, 80}

[0234] The reward function is set as equivalent fuel consumption and SOC maintenance. The equivalent fuel consumption refers to the amount of fuel required for the hybrid vehicle to operate in the current state. The smaller its value, the more energy-efficient the hybrid vehicle operates. And SOC maintenance refers to the degree of maintaining the remaining battery power. The larger its value, the longer the battery life. In this embodiment, the gradient descent method is used to update the parameters, and the network parameters are updated through backpropagation, as shown in the following formula:

[0235]

[0236] The Fifth Embodiment

[0237] In this embodiment, based on any of the foregoing embodiments, in order to study the effectiveness of the speed optimization method proposed in this embodiment, after using SimuLink to output the overtaking speeds of different driver styles, according to the designed speed optimization method, the speeds of the acceleration process, deceleration process, and constant speed process are optimized. The energy consumption of different-style drivers during overtaking is simulated through the DQN model. With the goal of minimizing the equivalent fuel consumption, the overtaking energy consumption before and after optimization is compared. As shown in FIGS. 13(a)-(c), they are respectively the SOC change curves and fuel consumption of a conservative driver style during overtaking before optimization, the SOC change curves and fuel consumption of a normal driver style during overtaking before optimization, and the SOC change curves and fuel consumption of an aggressive driver style during overtaking before optimization; FIGS. 14(a)-(c) correspond to the SOC change curves and fuel consumption of a conservative driver style during overtaking after optimization, the SOC change curves and fuel consumption of a normal driver style during overtaking after optimization, and the SOC change curves and fuel consumption of an aggressive driver style during overtaking after optimization.

[0238] At the same time, the SOC generated during the overtaking process of different driver styles is converted into equivalent fuel consumption, which is then superimposed on the fuel consumption lost by the vehicle engine during this process to obtain the total energy consumption of the overtaking process before and after speed optimization for different driver styles, as shown in Table 6:

[0239] Table 6. Total Overtaking Energy Consumption of Different Driver Styles Before and After Speed Optimization

[0240]

[0241] From Table 6, it is found that the equivalent fuel energy consumption during the overtaking process of conservative drivers before optimization is 3650 g, that of normal drivers is 6050 g, and that of aggressive drivers is 9700 g; the energy consumption of conservative drivers can be reduced by 9.59% after overtaking speed optimization, that of normal drivers can be reduced by 34.62%, and that of aggressive drivers can be reduced by 57.73%. The proposed optimization method can reduce the energy consumption by at least 9.59%.

[0242] Sixth Embodiment

[0243] In this embodiment, based on any of the foregoing embodiments, in order to verify the fuel economy before and after distinguishing driver styles, the speed planning output values in the previous chapter under different planning cycles are optimized and used as the input values of the DQN algorithm. The DQN model is used to simulate the energy consumption of different styles of drivers in the urban road overtaking environment to obtain the overtaking energy consumption of various driver styles after distinguishing driver styles, as Figure 15 shown. At the same time, the overtaking speed before distinguishing driver styles is also used as the input of the DQN model to simulate the energy consumption, and the SOC change curve and fuel consumption before distinguishing driver styles as shown in Figure 15 are obtained.

[0244] It can be seen that the equivalent fuel energy consumption of conservative driver styles during overtaking is 3300 g, that of normal driver styles is 3800 g, and that of aggressive driver styles is 4100 g. By calculating the overtaking energy consumption before distinguishing driver styles as shown in Figure 15 , the equivalent fuel energy consumption obtained is 6850 g. By comparison, it is found that the energy consumption generated during overtaking is significantly reduced after distinguishing driver styles.

[0245] In addition, those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the vehicle control system to implement the process steps of the embodiments of the above methods.

[0246] Therefore, the present application also provides a computer-readable storage medium storing a vehicle control program for an autonomous driving overtaking scenario that combines path planning and driving style. When the vehicle control program for the autonomous driving overtaking scenario that combines path planning and driving style is executed by a processor, it implements each step of the vehicle control method for the autonomous driving overtaking scenario that combines path planning and driving style as described in the above embodiments.

[0247] Among them, the computer-readable storage medium can be various computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.

[0248] It should be noted that since the storage medium provided in the embodiments of the present application is the storage medium used to implement the methods of the embodiments of the present application, those skilled in the art can understand the specific structure and variations of the storage medium based on the methods introduced in the embodiments of the present application, so it will not be elaborated here. Any storage medium used in the methods of the embodiments of the present application falls within the scope of protection of the present application.

[0249] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0250] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementation in the process Figure 1one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks

[0251] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks

[0252] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks

[0253] It should be noted that, in the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several means, several of these means may be embodied by one and the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names

[0254] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic inventive concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present application

[0255] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations

Claims

1. A vehicle control method for an automatic driving overtaking scenario combining path planning and driving style, characterized in that: Applied to a hybrid vehicle including a battery and an engine, the method comprises the following steps: Acquiring historical driving characteristic parameters of the vehicle, and determining a target driving style type according to the historical driving characteristic parameters, wherein the historical driving characteristic parameters include a lateral acceleration mean, a lateral velocity mean, a steering wheel angle, a lateral velocity standard deviation, a lateral acceleration standard deviation, and an average acceleration, and the target driving style includes an aggressive driving style, a normal driving style, and a conservative driving style; Determining a planned overtaking path and an overtaking speed that meet the target driving style based on a quadratic programming algorithm; Based on the planned overtaking path, the overtaking speed and the target driving style, an upper-layer speed optimization strategy and a lower-layer deep Q-learning energy management strategy are formulated; The speed constraint of the vehicle when overtaking is controlled to satisfy the upper-level speed optimization strategy, and the power distribution between the battery and the engine of the vehicle is controlled to satisfy the energy management strategy of the lower-level deep Q learning.

2. The method according to claim 1, characterized in that The step of determining the target driving style type according to the historical driving characteristic parameters comprises: Based on the principal component analysis method, the historical driving characteristic parameters are reduced in dimension to form a preprocessing data set; Calculate the Euclidean distance between each point in the preprocessed data set and the center point; Performing K-means clustering based on the Euclidean distance to calculate the lateral velocity center point value, the lateral acceleration center point value, and the steering wheel mean center point value; The driving style type corresponding to the interval in which the lateral velocity center point value, the lateral acceleration center point value and the steering wheel mean center point value are located is determined as the target driving style type.

3. The method according to claim 1, characterized in that The step of determining the planned overtaking path and overtaking speed that meets the target driving style based on the quadratic programming algorithm comprises: Planning an initial overtaking path that conforms to the target driving style based on a heuristic algorithm; The local path where the vehicle is located is intercepted in the initial overtaking path, the projection point of the vehicle is used as the coordinate origin, points within a suitable length are selected for smoothing, and the set of smoothed points is used as a reference line; Determining a cost function value of the reference line based on a preset line trajectory quality cost function; Iteratively calculate the overtaking path, and stop when the cost function value is less than the minimum value of the upper boundary of the local convex space, or greater than the maximum value of the lower boundary of the local convex space, to obtain the current optimal overtaking path, wherein the local convex space is a subspace of the complete constraint space composed of multiple local paths; The current optimal overtaking path obtained when the iteration stops is converted into Cartesian coordinates to obtain the overtaking planning path; An ST diagram is established based on the overtaking planning path, and based on the ST diagram, the overtaking speed of the vehicle is planned using dynamic programming, wherein S represents the driving distance of the vehicle along the overtaking planning path, and T represents the length of driving time.

4. The method according to claim 3, characterized in that The expression of the preset line trajectory quality cost function is: H(x)=w r f1(x,y)+w s f2(x,y)+w c f3(x,y) Where f1(x,y) represents the similarity cost function, f2(x,y) represents the smoothing cost function, f3(x,y) represents the compact cost function, w r is the weight of similar cost, w s is the weight of the smoothing cost; w c is the weight of the compactness cost.

5. The method according to claim 1, characterized in that The step of establishing an ST graph based on the planned overtaking path, and planning the overtaking speed of the vehicle by using dynamic programming based on the ST graph includes: Taking the planned overtaking path as the coordinate axis, establishing a Frenet coordinate system, projecting the preceding vehicle of the vehicle into the Frenet coordinate system, and establishing an ST graph, wherein S represents the driving distance of the vehicle along the planned overtaking path, and T represents the driving time length; Based on the ST diagram, the overtaking speed of the vehicle is planned by dynamic programming, wherein the dynamic programming satisfies the preceding vehicle distance constraint w o , speed constraint w v and the acceleration constraint w a ; The distance between the front of the vehicle and the rear of the preceding vehicle is taken as the preceding vehicle distance constraint w o ; The speed constraint w v The expression is: <h2 style=";text-align:left;direction:ltr">w<h2 style=";text-align:left;direction:ltr"> v <h2 style=";text-align:left;direction:ltr"> =f2(v-v0)<h2 style=";text-align:left;direction:ltr"> 2 The acceleration constraint w a for: In a =f3a 2 a=(v-v1)×Δt -1 Where v0 is the road speed limit, f2 is the speed cost weight, f3 is the acceleration cost weight, v is the current speed of the vehicle, v1 is the speed of the vehicle at the planning starting point, and Δt is the time difference between v and v1.

6. The method according to claim 1, characterized in that The upper layer speed optimization strategy includes acceleration strategy, constant speed strategy and deceleration strategy, among which: The acceleration process in the acceleration strategy satisfies the following expression: Where V(t, α) is the final speed of the vehicle during acceleration, V0 is the initial speed of the vehicle before acceleration; t0 and t e are the initial and final moments of the acceleration process, respectively, and the value range of α is (1, ∞); The vehicle speed in the constant speed strategy is the vehicle speed at which the fuel consumption is minimized; The deceleration process in the deceleration strategy satisfies the following expression: Where V is the final speed of the vehicle during deceleration, V0 is the initial speed of the vehicle during deceleration, γ≥1, t0 is the initial moment of the deceleration process, t e is the final moment of the deceleration process.

7. The method according to claim 1, characterized in that The energy management strategy of the lower-level deep Q learning includes state S, action A, reward R and time T, where: The state S includes the remaining charge in the battery and the required power of the engine, wherein the required power range is set to [-80, 80], and the remaining charge in the battery range is set to [0, 100]; The action A is the engine-battery power, and the action variable is discretized into 11 actions in the range of [0, 80]: A={0,10,20,30,40,50,60,70,80} The reward R is set to equivalent fuel consumption and SOC maintenance.

8. A vehicle control system, characterized in that: The vehicle control system includes: a memory, a processor, and a vehicle control program for an autonomous driving overtaking scenario combining path planning and driving style, which is stored in the memory and can be run on the processor. When the vehicle control program for an autonomous driving overtaking scenario combining path planning and driving style is executed by the processor, the steps of a vehicle control method for an autonomous driving overtaking scenario combining path planning and driving style are implemented as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a vehicle control program for an autonomous driving overtaking scenario that combines path planning and driving style. When the vehicle control program for an autonomous driving overtaking scenario that combines path planning and driving style is executed by a processor, the steps of a vehicle control method for an autonomous driving overtaking scenario that combines path planning and driving style are implemented as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Vehicle control methods, devices, electronic equipment and computer-readable media

    CN115571165B