End-to-end decision planning method for driving path of autonomous vehicle

Through end-to-end decision planning methods and Transformer neural network, the problems of driving safety and efficiency of autonomous vehicles in complex environments are solved, and more accurate and efficient driving trajectory planning is achieved.

CN120176675APending Publication Date: 2025-06-20HEFEI UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510326244.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing decision-making and planning methods for autonomous driving vehicles are difficult to ensure driving safety and efficiency when dealing with complex environments and uncertain driving behaviors.

Method used

Using an end-to-end decision planning method, the uncertainty of driving behavior is quantified through deep learning Transformer neural network, the time-space distance map is generated, the optimal target distance is calculated, and the driving trajectory is generated in the space-time safety corridor.

Benefits of technology

It improves the safety and efficiency of autonomous driving vehicles driving in complex environments, reduces the dependence on manual feature engineering and intermediate complex preprocessing, and enhances the prediction accuracy and overall performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120176675A_ABST
    Figure CN120176675A_ABST
Patent Text Reader

Abstract

The invention discloses an end-to-end decision planning method for a driving path of an automatic driving vehicle, and the method comprises the steps: calculating a candidate driving lane of the vehicle based on a road topological structure and navigation information of the vehicle; quantifying the driving behavior uncertainty using a risk assessment function that considers longitudinal and transverse collision risks and a vehicle shape; generating a discrete space-time interval distribution diagram based on the prediction information of the surrounding vehicles and the risk assessment result in combination with the lane boundary; establishing an automatic driving vehicle decision network, and training through a human driving data set to obtain an optimal target distance; and according to the optimal target distance, constructing a space-time safety corridor to ensure that the space-time safety corridor is not overlapped with surrounding vehicles, and generating a driving track. According to the method, the space-time interval diagram is generated by quantifying the uncertainty of the driving behavior, the optimal target interval is obtained by utilizing the neural network training based on transformer, and then the reliable safe track is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and is used for planning a driving trajectory during vehicle autonomous driving. Specifically, it relates to an end-to-end decision-making and planning method for the driving path of an autonomous vehicle. Background Art

[0002] In recent years, the automotive industry has witnessed a rapid development of intelligence, and the intelligence level of vehicles has been increasing day by day. Currently, the sales proportion of vehicles with autonomous driving capabilities at L2 level and above has approached 40%. As autonomous driving technology gradually moves towards popularization, safety has always been the core point that attracts much attention. In actual driving scenarios, autonomous vehicles often encounter various complex situations, such as vehicles suddenly changing lanes forcefully, or non-motor vehicles with blocked vision making it difficult to predict the action trajectory, which makes ensuring driving safety a thorny problem in scientific research.

[0003] With the gradual in-depth research, autonomous driving has been realized in complex environments full of other traffic participants and obstacles. However, in order to ensure safety, existing decision-making and planning methods often sacrifice some driving efficiency. In addition to considering driving efficiency, in the actual driving environment, the behaviors of surrounding vehicles are often unpredictable, such as sudden hard braking, forced lane change or sudden failure, etc. For these problems affecting driving safety, the existing decision-making and planning methods are still difficult to guarantee the accuracy of environmental information processing and planning, which makes the safety of autonomous vehicles during driving an urgent problem to be solved. Summary of the Invention

[0004] Aiming at the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to provide an end-to-end decision-making and planning method for the driving path of an autonomous vehicle, so as to solve the problems in the existing technologies that vehicle behavior planning is not timely and the planned path is greatly affected by environmental factors. The present invention quantifies the uncertainty of driving behaviors, generates a spatio-temporal distance map, and uses a neural network based on transformer to train to obtain the optimal target distance, and then generates a reliable safe trajectory.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] An end-to-end decision-making and planning method for the driving path of an autonomous vehicle of the present invention comprises the following steps:

[0007] 1) Based on the road topology structure and the ego-vehicle navigation information, calculate the candidate lanes for the ego-vehicle to drive through the depth-first search algorithm;

[0008] 2) Use a risk assessment function considering longitudinal and lateral collision risks and vehicle shape to quantify the uncertainty of driving behaviors;

[0009] 3) Based on the prediction information of surrounding vehicles and the results of risk assessment, combined with lane boundaries, a discrete time-space distance distribution map is generated;

[0010] 4) Based on the lane information, vehicle historical trajectory information and vehicle navigation information, an autonomous driving vehicle decision network is established based on the Transformer neural network architecture, and the optimal target spacing is obtained through training using human driving data sets;

[0011] 5) Based on the optimal target spacing, a spatiotemporal safety corridor is constructed to ensure that there is no overlap with surrounding vehicles, and a driving trajectory is generated within the spatiotemporal safety corridor based on the IDM-DP trajectory generation method and quadratic programming algorithm.

[0012] Furthermore, the road topology structure in step 1) is obtained through an internal map data set, and the structure includes the connection relationship of the roads, the number of lanes and turning restrictions.

[0013] Furthermore, in step 1), the navigation information of the vehicle is obtained through an internal map data set, including the current position of the vehicle, the destination and the planned driving route.

[0014] Furthermore, the candidate lane in step 1) refers to a lane option that may be changed or used in the trajectory planning method, and the specific generation process is: using a depth-first search algorithm, starting from the lane where the vehicle is currently located as the starting node, searching according to the connection relationship between lanes in the road topology structure according to the depth-first principle; in the search process, combined with the navigation information of the vehicle, after traversing the relevant lane nodes through the depth-first search algorithm, the lane that meets the conditions is determined as the candidate lane for the vehicle to travel.

[0015] Furthermore, in step 2), a distance vector is established that takes into account the longitudinal and lateral directions as well as the shape of the vehicle, and then the lateral and longitudinal safety distances are calculated according to the motion state of the vehicle, and the risk assessment function is obtained by combining them;

[0016] The distance vector between the vehicle and the surrounding vehicles is expressed as: The longitudinal component r of the distance vector along the vehicle's velocity si , transverse component r li for:

[0017]

[0018] Where s represents the longitudinal position of the vehicle in the Frenet coordinate system, l represents the lateral position of the vehicle in the Frenet coordinate system, and θ iDenote the heading angle of the vehicle corresponding to the number \(i\), where \(i\) represents the vehicle ID, \(i = 0, 1, 2, \ldots, n\). When \(i = 0\), it represents the ego vehicle, and when \(i = 1, 2, \ldots, n\), it represents the surrounding vehicles; \(s_0\) represents the longitudinal position of the ego vehicle in the Frenet coordinate system, and \(l_0\) represents the lateral position of the ego vehicle in the Frenet coordinate system; \(s\) i represents the longitudinal position of the surrounding vehicle in the Frenet coordinate system, and \(l\) i represents the lateral position of the surrounding vehicle in the Frenet coordinate system, \(i = 1, 2, \ldots, n\);

[0019] When considering the vehicle shape, the boundary of the vehicle is represented by the four corner points of the vehicle; according to the position relationship between the surrounding vehicle and the ego vehicle, the longitudinal and lateral distances between the four corner points of the ego vehicle and the surrounding vehicle are obtained, and the expression is as follows:

[0020]

[0021] In the formula, \(r\) represents the longitudinal and lateral distances between the four corner points of the ego vehicle and the surrounding vehicle, \(a\) and \(b\) respectively represent the length and width of the surrounding vehicle; \(w\) and \(l\) respectively represent the length and width of the ego vehicle; the \(poslin()\) function is a positive linear function; \(\pm\) represents the position relationship between the ego vehicle and the surrounding vehicle; \(S=\{[1,1] T ,[-1,1] T ,[-1,-1] T ,[1,-1] T \}; \(R\) is a rotation matrix, represents the yaw angle deviation between the ego vehicle and the surrounding vehicle;

[0022] Calculate the corresponding safety distances for the longitudinal and lateral directions respectively. Through the longitudinal and lateral safety distances and the distance vector between the vehicles, accurately evaluate the risk level of the vehicle;

[0023] The longitudinal safety distance \(D\) si is expressed as:

[0024]

[0025] In the formula, \(v_0\) represents the speed of the ego vehicle; \(v\) i represents the speed of the surrounding vehicle; \(sign()\) is a sign function; \(\zeta\) represents the reaction time, including the reaction time of the decision-making system and the time when the brake takes effect; \(G\) represents the expected distance between the ego vehicle and the surrounding vehicle; \(a\) min represents the minimum deceleration of the vehicle; \(\theta\) i represents the heading angle of the vehicle; \(b_1 = v\) i \(\zeta+G\), the lateral safety distance \(D\) li is expressed as:

[0026]

[0027] Hazard assessment function F i It is expressed as:

[0028]

[0029] The result of this hazard assessment function is affected by the speeds and heading angle information of the host vehicle and surrounding vehicles, forming a hazard field, representing the behavioral uncertainty of the vehicle in the form of a hazard field, and correcting the hazard area generated by the vehicle according to the maximum acceptable hazard level as the final hazard domain.

[0030] Furthermore, step 3) specifically includes:

[0031] Using all the candidate lanes obtained in step 1), after determining the candidate lanes according to the behavior of the host vehicle, predict the future trajectories of all vehicles on the candidate lanes to obtain the position and speed information of all vehicles at several future discrete moments; through the hazard assessment function in step 2), obtain the hazard assessment results at the corresponding moments, and combine the boundaries of each lane to obtain a discrete spatio-temporal distance distribution map.

[0032] Furthermore, in step 4), build an autonomous driving vehicle decision network based on the Transformer neural network and combine the tree search strategy to train and obtain the optimal target distance; specifically include:

[0033] 41) Encode the input data;

[0034] The input data representation: Receive input data based on the Transformer neural network architecture, including: the historical trajectory of the vehicle and environmental information; for each vehicle, its historical trajectory is a series of dynamic features in the past time steps T = 20, including the two-dimensional position, heading angle, speed, and boundary size of the vehicle; for environmental information, consider two types of map elements, including lanes and crosswalks, and each map element is represented by a vector, which consists of a series of nodes with different associated features; for each vehicle, construct a local environmental information; the position attributes of all vehicles and map elements are converted to the local coordinate system of the AV;

[0035] Lane line information encoding: Discretize the lane lines, sample at a certain interval, and each sampling point is represented by coordinates (x, y), and at the same time append features, including line type and curvature; normalize the features and splice them into a multi-dimensional vector; use position encoding to add position information to each sampling point, and use sine-cosine position encoding, and the encoding formula is as follows:

[0036]

[0037] Among them, PE is the abbreviation of Positional Encoding. pos represents the position in the sequence. For the even - numbered dimensions, the dimension value is 2i and the sine function is used; for the odd - numbered dimensions, the dimension value is 2i + 1 and the cosine function is used. d represents the total number of dimensions;

[0038] Encode the lane lines in a sequence form to meet the input requirements of the Transformer neural network architecture;

[0039] Encoding of vehicle historical trajectory information: The content recorded by the vehicle historical trajectory in time series includes vehicle position, speed, and acceleration. Concatenate the data at different times into a vector in sequence. For the time steps at t1, t2,..., t n (a total of T time steps), the corresponding positions (x1, y1), (x2, y2),...,(x n ,y n ), speed v1, v2,..., v n are combined into a long vector. Add positional encoding to make the vector have a chronological order and convert it into a sequence with position awareness that can be processed by the Transformer neural network architecture;

[0040] Encoding of environmental information: The environmental information encoder consists of a lane encoder for processing lanes and a crosswalk encoder for processing crosswalk vectors. The lane encoder uses a multi - layer perceptron to encode numerical features and an embedding layer to encode discrete features. The crosswalk encoder is another multi - layer perceptron for encoding numerical features;

[0041] 42) Select the query, key, and value;

[0042] Selection of query Q:

[0043] After encoding the position coordinates, speed, and acceleration information of the vehicle in the past multiple time steps, use them as part of Q. Embed and represent the environmental information related to the current position of the vehicle, including the curvature of the surrounding roads, the status of traffic signs and signals, and the relative positions of other vehicles, and also use them as part of Q, so that the Transformer neural network can consider the impact of environmental factors on the vehicle trajectory during prediction;

[0044] Selection of key K:

[0045] Use the detailed information in the vehicle historical trajectory as K, including the change in the steering angle of the vehicle and the driving directions at different times. Further encode the environmental information as K, including the topological structure information of the road, the positions and types of surrounding buildings;

[0046] Selection of value V:

[0047] The value V of the vehicle historical trajectory includes the average value of the vehicle speed, the standard deviation of the acceleration, and the short-term trend information predicted based on the vehicle historical trajectory within a past period of time;

[0048] The value V of the environmental information is the result of fusing and abstracting environmental characteristics, including aggregating and representing the traffic flow conditions within a certain range around the vehicle;

[0049] Taking the query, key, and value as inputs to calculate the output of the attention mechanism as follows:

[0050]

[0051] where d k is the dimension of the key K, and K T represents the transpose of the K matrix;

[0052] Passing the result through softmax to normalize the scores of all input quantities, so that the obtained scores are all positive and sum to 1;

[0053] 43) Encoding process;

[0054] Embedding layer: For the vehicle historical trajectory, use the linear layer in the Transformer neural network to map it to a high-dimensional space to obtain a matrix of T×d, where d is the embedding dimension; for the environmental information, use the linear layer to map it to a d-dimensional vector;

[0055] Multi-head attention layer: Input the Q, K, and V generated by the encoding sequence into the multi-head attention module built into the multi-head attention layer of the Transformer neural network architecture; multiple heads calculate the attention in parallel, and different heads focus on different subspace features of the input to capture the complex shapes of lane lines, the fluctuations of trajectory speeds, and the detours of navigation routes;

[0056] Feed-forward network layer: Input the output result of the multi-head attention layer into the feed-forward neural network built into the feed-forward network layer. The feed-forward neural network is a structure containing two fully connected layers and activation functions, which is used to perform non-linear transformation on features, amplify important decision signals, suppress noise, and make the lane lines, vehicle driving trajectories, and self-vehicle navigation information more distinguishable for decision-making; the expression of the feed-forward neural network is: FFN(x) = max(0, xW1 + b1)W2 + b2, where W1, W2 are weight matrices, b1, b2 are bias vectors, and FFN represents the feed-forward neural network;

[0057] 44) Decoding process;

[0058] Masked Multi-Head Self-Attention Layer: The masked multi-head self-attention layer is used to ensure that the neural network only focuses on previous time steps when generating predictions for each time step. Specifically, a mask M (lower triangular matrix) is created, where the elements on and below the diagonal are 0, and the elements above the diagonal are negative infinity. The mask matrix M is a 20×20 lower triangular matrix, with elements on and below the diagonal being 0 and elements above the diagonal being negative infinity.

[0059] The mask matrix M is added to the attention score matrix before softmax processing. When calculating softmax, the elements above the diagonal (corresponding to information from future time steps) will approach 0, thus ensuring that the model is not affected by future information when generating the output for the current time step.

[0060] Cross-Attention Layer: Attention calculation is performed between the output of the decoder and the output of the encoder to fuse the information from the encoding part. The required Q matrix comes from the output of the decoder, and the K matrix and V matrix come from the output of the encoder.

[0061] Feed-Forward Neural Network: The output of the encoder-decoder attention layer passes through a feed-forward neural network to obtain the final output of the decoder.

[0062] Furthermore, the optimal target spacing in step 4) is calculated by comprehensively considering the cost magnitude, where the cost magnitude includes the arrival cost and efficiency cost of calculating each spacing. Among them:

[0063] Arrival Cost C arrival includes overlap cost C overlap and expansion cost C extend ;

[0064] Overlap Cost C overlap is expressed as:

[0065]

[0066] Among them,

[0067]

[0068] Δv = v s - v e

[0069] where τ is the time corresponding to a lane-changing behavior, v s represents the speed of surrounding vehicles, and v e represents the speed of the ego vehicle;

[0070] Expansion Cost C extend The expression is:

[0071]

[0072] Among them, represents the front-end position of the spacing where the vehicle is located, represents the rear-end position of the spacing where the vehicle is located, represents the length of the spacing where the vehicle is located, represents the front-end position of the target spacing, represents the rear-end position of the target spacing;

[0073] Arrival cost C arrivall The expression is:

[0074] C arrivall = C avarlap + C extend ;

[0075] Efficiency cost C eff The expression is:

[0076]

[0077] In the formula, V lim it represents the speed limit of the lane where the spacing is located, V front represents the speed magnitude of the vehicle in front of the spacing, V rear represents the speed magnitude of the vehicle behind the spacing.

[0078] Furthermore, in step 4), the optimal target spacing uses a tree search strategy to judge the transfer relationship between spacings. Specifically, it includes taking the current spacing where the vehicle is located as the root node of the tree, and the spacing at each moment corresponds to the nodes on each layer of the search tree. If there is a transfer relationship between two spacings, the corresponding two nodes are connected. At this time, the spacing where the other node except the root node is located is used as the optimal target spacing; the Transformer neural network verifies according to the target spacing sequence using the existing human driving data set, and then outputs the optimal target spacing.

[0079] Furthermore, step 5) specifically includes:

[0080] Based on the discrete spatio-temporal spacing distribution map obtained in step 3), according to the spacing sequence in the discrete spatio-temporal spacing distribution map, by interpolating the boundaries of the spacing, a spatio-temporal safety corridor is obtained.

[0081] Furthermore, the solution process of the driving trajectory in step 5) includes:

[0082] According to the state of the vehicle ahead at the target spacing, a corresponding longitudinal point set is generated based on the Intelligent Driver Model (IDM), and then a corresponding lateral point set is generated according to the lateral width of the spatio-temporal safety corridor. The longitudinal point set and the lateral point set are combined to form a point set. The number of points in the point set is p = 1, 2, 3... n, and each point in the point set has a longitudinal coordinate s p and a lateral coordinate l p ; During the dynamic programming (DP) process, three costs are considered to ensure safety and comfort, including the smoothness cost C smooth , the cost C bounds of being restricted within the corridor, and the centering cost C middle ; The cost function C total in the dynamic programming process is as follows:

[0083] C smooth = abs((l p - l p-1 ) / (s p - s p-1 ))

[0084]

[0085] C middle = (l max + l min ) / 2

[0086]

[0087] where p represents the points in the point set, s p , s p-1 represent the longitudinal coordinates of the p-th point and the (p - 1)-th point, and l p , l p-1 represent the lateral coordinates of the p-th point and the (p - 1)-th point. l max , l min represent the maximum and minimum values of the corresponding time boundaries respectively; inf represents infinity, and abs represents solving the absolute value; ω smooth , ω bounds , ω middle are the weights of the corresponding terms;

[0088] Using the quadratic programming algorithm, the reference trajectory is optimized. In the process of establishing the optimization equation, that is, the objective function, the method of minimizing jerk is adopted; The optimization variables are selected where χ represents the optimization variable, s represents the longitudinal coordinate of the sampling point, l represents the lateral coordinate of the sampling point, represent the first-order, second-order, and third-order derivatives of the longitudinal coordinate with respect to time respectively, respectively represent the first-order, second-order, and third-order derivatives of the horizontal coordinate with respect to time;

[0089] The objective function consists of the closeness to the reference trajectory and comfort, where the closeness to the reference trajectory is expressed as The comfort cost includes the speed cost acceleration cost and jerk cost

[0090] The complete objective function is expressed as:

[0091]

[0092] In the formula, g represents the reference trajectory; p represents the reference point; n represents the total number of sampling points; χ i , χ i ', χ i ”, χ i ”' respectively represent the p-th optimization variable and the first-order, second-order, and third-order derivatives of the p-th optimization variable, ω χ , ω χ' , ω χ” , ω χ”' respectively represent the weights of the closeness, speed cost, acceleration cost, and comfort cost;

[0093] The optimization equation is organized into a quadratic programming problem and solved using the QSQP solver to obtain the final driving trajectory.

[0094] Advantages of the present invention:

[0095] The neural network based on Transformer constructed by the present invention adopts an end-to-end architecture, solving the problem of predicting the future driving trajectory of vehicles; this neural network directly uses vehicle historical trajectory data and environmental information as inputs and outputs the optimal target spacing, without relying on any manual feature engineering or complex preprocessing in the intermediate stage.

[0096] By virtue of the sequence modeling and long-distance dependence capturing capabilities of the Transformer neural network, the neural network architecture automatically learns the hidden spatio-temporal features and complex correlations in the input data during the training process; the multi-head self-attention mechanism of Transformer enables the model to adaptively focus on the key parts of the historical trajectory and environmental information, discovering potential patterns in the data.

[0097] The present invention does not require manual setting of intermediate targets or step-by-step processing, and directly maps the original input data to the output of the vehicle's possible future driving trajectory. This end-to-end design avoids the error accumulation problem caused by the series connection of multiple independent modules in the traditional method, ensuring the accuracy of the prediction and the overall performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] Figure 1 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION

[0099] In order to facilitate the understanding of those skilled in the art, the present invention is further described below in conjunction with embodiments and drawings. The contents mentioned in the implementation modes are not intended to limit the present invention.

[0100] Reference Figure 1 As shown, an end-to-end decision-making planning method for an autonomous driving vehicle driving path of the present invention comprises the following steps:

[0101] 1) Based on the road topology and the navigation information of the ego vehicle, the candidate lanes for the ego vehicle are calculated through a depth-first search algorithm;

[0102] The road topology structure in step 1) is obtained through an internal map data set, and the structure includes the connection relationship of the roads, the number of lanes and turning restrictions.

[0103] Wherein, in said step 1), the navigation information of the vehicle is obtained through an internal map data set, including the current position of the vehicle, the destination and the planned driving route.

[0104] The candidate lane in step 1) refers to a lane option that may be changed or used in the trajectory planning method, and the specific generation process is: using a depth-first search algorithm, starting from the lane where the vehicle is currently located as the starting node, searching according to the connection relationship between lanes in the road topology structure according to the depth-first principle; in the search process, combined with the navigation information of the vehicle, after traversing the relevant lane nodes through the depth-first search algorithm, the lane that meets the conditions is determined as the candidate lane for the vehicle to travel.

[0105] 2) Quantify driving behavior uncertainty using a risk assessment function that considers longitudinal and lateral collision risks and vehicle shape;

[0106] Wherein, in the step 2), a distance vector is established taking into account the longitudinal and lateral distances as well as the shape of the vehicle, and then the lateral and longitudinal safety distances are calculated according to the motion state of the vehicle, and the risk assessment function is obtained after combining them;

[0107] The distance vector between the vehicle and the surrounding vehicles is expressed as: The longitudinal component r of the distance vector along the vehicle's velocity si , transverse component rli is:

[0108]

[0109] where s represents the longitudinal position of the vehicle in the Frenet coordinate system, l represents the lateral position of the vehicle in the Frenet coordinate system, and θ i represents the heading angle of the vehicle corresponding to the vehicle numbered i, i represents the id of the vehicle, i = 0, 1, 2, …, n. When i = 0, it represents the ego vehicle, and when i = 1, 2, …, n, it represents the surrounding vehicles; s0 represents the longitudinal position of the ego vehicle in the Frenet coordinate system, and l0 represents the lateral position of the ego vehicle in the Frenet coordinate system; s i represents the longitudinal position of the surrounding vehicle in the Frenet coordinate system, and l i represents the lateral position of the surrounding vehicle in the Frenet coordinate system, i = 1, 2, …, n;

[0110] When considering the vehicle shape, the boundary of the vehicle is represented by the four corner points of the vehicle; according to the position relationship between the surrounding vehicle and the ego vehicle, the longitudinal and lateral distances between the four corner points of the ego vehicle and the surrounding vehicle are obtained, and the expression is as follows:

[0111]

[0112] In the formula, r represents the longitudinal and lateral distances between the four corner points of the ego vehicle and the surrounding vehicle, a and b respectively represent the length and width of the surrounding vehicle; w and l respectively represent the length and width of the ego vehicle; the poslin() function is a positive linear function; ± represents the position relationship between the ego vehicle and the surrounding vehicle; S = {[1, 1] T , [-1, 1] T , [-1, -1] T , [1, -1] T}; R is the rotation matrix, represents the yaw angle deviation between the ego vehicle and the surrounding vehicle;

[0113] The corresponding safety distances are calculated separately for the longitudinal and lateral directions, and through the longitudinal and lateral safety distances and the distance vector between the vehicles, the risk level of the vehicle is accurately evaluated;

[0114] The longitudinal safety distance D si is expressed as:

[0115]

[0116] In the formula, v0 represents the speed of the ego vehicle; v i represents the speed of the surrounding vehicle; sign() is the sign function; ζ represents the reaction time, including the reaction time of the decision-making system and the time when the brake takes effect; G represents the expected distance between the ego vehicle and the surrounding vehicle; amin represents the minimum deceleration of the vehicle; θ i represents the heading angle of the vehicle; b1 = v i ζ + G, the lateral safety distance D li is expressed as:

[0117]

[0118] The risk assessment function F i is expressed as:

[0119]

[0120] The result of the risk assessment function is affected by the speeds and heading angle information of the host vehicle and surrounding vehicles, forming a risk field, representing the behavioral uncertainty of the vehicle in the form of a risk field, and correcting the risk area generated by the vehicle according to the maximum acceptable risk as the final risk domain.

[0121] 3) Based on the prediction information of surrounding vehicles and the risk assessment results, combined with the lane boundaries, generate a discrete time-space gap distribution map; specifically including:

[0122] Using all the candidate lanes obtained in step 1), after determining the candidate lanes according to the behavior of the host vehicle, predict the future trajectories of all vehicles on the candidate lanes to obtain the position and speed information of all vehicles at several future discrete moments; through the risk assessment function in step 2), obtain the risk assessment results at the corresponding moments, and combined with the boundaries of each lane, obtain a discrete time-space gap distribution map.

[0123] 4) Based on the lane line information, vehicle historical trajectory information and host vehicle navigation information, establish an autonomous driving vehicle decision-making network based on the Transformer neural network architecture, and train to obtain the optimal target spacing through a human driving data set;

[0124] Among them, in step 4), an autonomous driving vehicle decision-making network is built based on the Transformer neural network and combined with a tree search strategy to train and obtain the optimal target spacing; specifically including:

[0125] 41) Encode the input data;

[0126] Input data representation: Receive input data based on the Transformer neural network architecture, including: vehicle historical trajectories and environmental information; for each vehicle, its historical trajectory is a series of dynamic features in the past time steps T = 20, including the two-dimensional position, heading angle, speed, and boundary size of the vehicle; for environmental information, two types of map elements are considered, including lanes and crosswalks, and each map element is represented by a vector, which consists of a series of nodes with different associated features; for each vehicle, construct a local environmental information; the position attributes of all vehicles and map elements are converted to the local coordinate system of the AV.

[0127] Lane line information encoding: Discretize the lane lines, sample at a certain interval, and each sampling point is represented by coordinates (x, y), and additional features are included, including line type and curvature; normalize the features and splice them into a multi-dimensional vector; use positional encoding to add position information to each sampling point, and use sine-cosine positional encoding. The encoding formula is as follows:

[0128]

[0129] where PE is the abbreviation of Positional Encoding, pos represents the position in the sequence, the dimensions of the even positions use 2i and the sin function, the dimensions of the odd positions use 2i + 1 and the cos function, and d represents the total number of dimensions;

[0130] Encode the lane lines in a sequence form to adapt to the input requirements of the Transformer neural network architecture.

[0131] Vehicle historical trajectory information encoding: The content recorded in the vehicle historical trajectory in time series includes vehicle position, speed, and acceleration; splice the data at different times into a vector in sequence, and for the time steps at t1, t2,..., t n (a total of T time steps) corresponding positions (x1, y1), (x2, y2),...,(x n , y n ), speed v1, v2,..., v n Combine them into a long vector; add positional encoding to make the vector have a chronological order and convert it into a sequence with position awareness that can be processed by the Transformer neural network architecture.

[0132] Environmental information encoding: The environmental information encoder consists of a lane encoder for processing lanes and a crosswalk encoder for processing crosswalk vectors; the lane encoder uses a multi-layer perceptron (MLP) to encode numerical features and an embedding layer to encode discrete features; the crosswalk encoder is another multi-layer perceptron for encoding numerical features.

[0133] 42) Select the query, key, and value;

[0134] Selection of query Q:

[0135] Encode the position coordinates, speed, and acceleration information of the vehicle at multiple past time steps and use them as part of Q; Embed the environmental information related to the current position of the vehicle, including the curvature of the surrounding roads, the status of traffic signs and signals, and the relative positions of other vehicles, and also use it as part of Q, so that the Transformer neural network takes into account the impact of environmental factors on the vehicle trajectory during prediction;

[0136] Selection of key K:

[0137] Use the detailed information in the vehicle's historical trajectory as K, including the change in steering angle and the driving direction at different times; Further encode the environmental information as K, including the topological structure information of the road and the positions and types of surrounding buildings;

[0138] Selection of value V:

[0139] The value V of the vehicle's historical trajectory includes the average value of the vehicle speed, the standard deviation of the acceleration, and the short-term trend information predicted based on the vehicle's historical trajectory over a past period of time;

[0140] The value V of the environmental information is the result of fusing and abstracting environmental features, including aggregating and representing the traffic flow situation within a certain range around;

[0141] Use the query, key, and value as inputs to calculate the output of the attention mechanism as follows:

[0142]

[0143] where d k is the dimension of the key K, and K T represents the transpose of the K matrix;

[0144] Pass the result through softmax to normalize the scores of all input quantities, so that the obtained scores are all positive and sum to 1;

[0145] 43) Encoding process;

[0146] Embedding layer: For the vehicle's historical trajectory, use the linear layer in the Transformer neural network to map it to a high-dimensional space to obtain a matrix of T×d, where d is the embedding dimension; For the environmental information, use the linear layer to map it to a d-dimensional vector;

[0147] Multi-Head Attention Layer: Input the Q, K, and V generated from the encoded sequence into the multi-head attention module built into the multi-head attention layer of the Transformer neural network architecture; multiple heads calculate attention in parallel, and different heads focus on different subspace features of the input to capture the complex shapes of lane lines, fluctuations in trajectory speed, and detours in navigation routes;

[0148] Feed-Forward Network Layer: Input the output result of the multi-head attention layer into the feed-forward neural network built into the feed-forward network layer. The feed-forward neural network is a structure that includes two fully connected layers and an activation function, which is used to perform non-linear transformation on features, amplify important decision signals, suppress noise, and make the lane lines, vehicle driving trajectories, and self-vehicle navigation information more distinguishable for decision-making; The expression of the feed-forward neural network is: FFN(x) = max(0, xW1 + b1)W2 + b2, where W1 and W2 are weight matrices, and b1 and b2 are bias vectors, and FFN represents the feed-forward neural network;

[0149] 44) Decoding process;

[0150] Masked Multi-Head Self-Attention Layer: Use the masked multi-head self-attention layer to ensure that the neural network only focuses on previous time steps when generating predictions for each time step; specifically: create a mask M (lower triangular matrix), where the elements on and below the diagonal are 0, and the elements above the diagonal are negative infinity; The mask matrix M is a 20×20 lower triangular matrix, with elements on and below the diagonal being 0 and elements above the diagonal being negative infinity;

[0151] Add the mask matrix M to the attention score matrix before softmax processing. When calculating softmax, the elements above the diagonal (corresponding to information from future time steps) will approach 0, thus ensuring that the model will not be affected by future information when generating the output for the current time step;

[0152] Cross-Attention Layer: Perform attention calculation between the output of the decoder and the output of the encoder to fuse the information in the encoded part, where the required Q matrix comes from the output of the decoder, and the K matrix and V matrix come from the output of the encoder;

[0153] Feed-Forward Neural Network: The output of the encoder-decoder attention layer passes through a feed-forward neural network to obtain the final output of the decoder.

[0154] Among them, the optimal target spacing in step 4) is calculated by comprehensively considering the cost size, and the cost size includes the arrival cost and efficiency cost of calculating each spacing, where:

[0155] Arrival cost C arrival includes overlap cost C overlap and expansion cost C extend ;

[0156] Overlap cost C overlap is expressed as:

[0157]

[0158] wherein,

[0159]

[0160] Δv = v s - v e

[0161] where τ is the time corresponding to a lane-changing behavior, v s represents the speed of surrounding vehicles, and v e represents the speed of the host vehicle;

[0162] Expansion cost C extend The expression is:

[0163]

[0164] wherein, represents the front position of the interval where the host vehicle is located, represents the rear position of the interval where the host vehicle is located, represents the length of the interval where the host vehicle is located, represents the front position of the target interval, represents the rear position of the target interval;

[0165] Arrival cost C arrivall The expression is:

[0166] C arrivall = C avarlap + C extend ;

[0167] Efficiency cost C eff The expression is:

[0168]

[0169] In the formula, V lim it represents the speed limit of the lane where the interval is located, V front represents the speed magnitude of the vehicle in front of the interval, V rear represents the speed magnitude of the vehicle behind the interval.

[0170] Specifically, in step 4), the optimal target spacing uses a tree search strategy to judge the transition relationship between spacings. Specifically, it includes taking the current spacing where the vehicle itself is located as the root node of the tree. The spacing at each moment corresponds to the nodes on each layer of the search tree. If there is a transition relationship between two spacings, the corresponding two nodes are connected. At this time, the spacing where the other node except the root node is located is taken as the optimal target spacing; the Transformer neural network validates according to the target spacing sequence using the existing human driving data set, and then outputs the optimal target spacing.

[0171] 5) Construct a spatio-temporal safety corridor according to the optimal target spacing to ensure non-overlap with surrounding vehicles, and generate a driving trajectory within the spatio-temporal safety corridor based on the IDM-DP trajectory generation method and the quadratic programming algorithm; specifically including:

[0172] Based on the discrete spatio-temporal spacing distribution map obtained in step 3), according to the spacing sequence in the discrete spatio-temporal spacing distribution map, by interpolating the boundaries of the spacings, a spatio-temporal safety corridor is obtained.

[0173] Among them, the solution process of the driving trajectory in step 5) includes:

[0174] According to the state of the vehicle in front of the target spacing, a corresponding longitudinal point set is generated based on the intelligent driver model (IDM), and then a corresponding lateral point set is generated according to the lateral width of the spatio-temporal safety corridor. The longitudinal point set and the lateral point set are combined together to form a point set. The number of points in the point set is p = 1, 2, 3... n. Each point in the point set has a longitudinal coordinate s p and a lateral coordinate l p ; in the dynamic programming (DP) process, three costs are considered to ensure safety and comfort, including the smoothing cost C smooth , the cost C bounds of being restricted within the corridor and the centering cost C middle ; the cost function C total in the dynamic programming process is:

[0175] C smooth = abs((l p - l p-1 ) / (s p - s p-1 ))

[0176]

[0177] C middle = (l max + l min ) / 2

[0178]

[0179] Among them, p represents the points in the point set, and s p , s p-1 represents the longitudinal coordinates of the p-th point and the (p - 1)-th point, and l p , l p-1 represents the transverse coordinates of the p-th point and the (p - 1)-th point, and l max 、l min respectively represent the maximum and minimum values of the boundary at the corresponding moment; inf represents infinity, and abs represents solving the absolute value; ω smooth 、ω bounds 、ω middle are the weights of the corresponding terms;

[0180] Using the quadratic programming algorithm, the reference trajectory is optimized. In the process of establishing the optimization equation, that is, the objective function, the method of minimize jerk is adopted; the optimization variables are selected Among them, χ represents the optimization variable, s represents the longitudinal coordinate of the sampling point, and l represents the transverse coordinate of the sampling point, respectively represent the first-order derivative, second-order derivative, and third-order derivative of the longitudinal coordinate with respect to time, respectively represent the first-order derivative, second-order derivative, and third-order derivative of the transverse coordinate with respect to time;

[0181] The objective function consists of the closeness to the reference trajectory and comfort. Among them, the closeness to the reference trajectory is expressed as The comfort cost includes the speed cost the acceleration cost and the jerk cost

[0182] The complete objective function is expressed as:

[0183]

[0184] In the formula, g represents the reference trajectory; p represents the reference point; n represents the total number of sampling points; χ i , χ i ', χ i ”, χ i ”' respectively represent the p-th optimization variable and the first-order, second-order, and third-order derivatives of the p-th optimization variable, ω χ , ω χ' , ω χ” , ω χ”' respectively represent the weights of the closeness, speed cost, acceleration cost, and comfort cost;

[0185] The optimization equation is arranged into a quadratic programming problem and solved using the QSQP solver to obtain the final driving trajectory.

[0186] The specific application scenarios of the present invention are numerous. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, several improvements can be made without departing from the principle of the present invention, and these improvements should also be regarded as within the protection scope of the present invention.

Claims

1. An end-to-end decision-making planning method for an autonomous driving vehicle driving path, characterized in that: Here are the steps: 1) Based on the road topology and the navigation information of the ego vehicle, the candidate lanes for the ego vehicle are calculated through a depth-first search algorithm; 2) Quantify driving behavior uncertainty using a risk assessment function that considers longitudinal and lateral collision risks and vehicle shape; 3) Based on the prediction information of surrounding vehicles and the results of risk assessment, combined with lane boundaries, a discrete time-space distance distribution map is generated; 4) Based on the lane information, vehicle historical trajectory information and vehicle navigation information, an autonomous driving vehicle decision network is established based on the Transformer neural network architecture, and the optimal target spacing is obtained through training using human driving data sets; 5) Based on the optimal target spacing, a spatiotemporal safety corridor is constructed to ensure that there is no overlap with surrounding vehicles, and a driving trajectory is generated within the spatiotemporal safety corridor based on the IDM-DP trajectory generation method and quadratic programming algorithm.

2. The end-to-end decision-making planning method for the driving path of an autonomous driving vehicle according to claim 1, characterized in that: The road topology structure in step 1) is obtained through an internal map data set, and the structure includes the connection relationship of the roads, the number of lanes and turning restrictions.

3. The end-to-end decision-making planning method for the driving path of an autonomous driving vehicle according to claim 1, characterized in that: In step 1), the vehicle navigation information is obtained through an internal map data set, including the vehicle's current location, destination, and planned driving route.

4. The end-to-end decision-making planning method for the driving path of an autonomous driving vehicle according to claim 1, characterized in that: The candidate lane in step 1) refers to a lane option that may be changed or used in the trajectory planning method. The specific generation process is: using a depth-first search algorithm, starting from the lane where the vehicle is currently located as the starting node, searching according to the connection relationship between lanes in the road topology structure according to the depth-first principle; in the search process, combined with the navigation information of the vehicle, after traversing the relevant lane nodes through the depth-first search algorithm, the lane that meets the conditions is determined as the candidate lane for the vehicle to travel.

5. The end-to-end decision-making planning method for the driving path of an autonomous driving vehicle according to claim 1, characterized in that: In the step 2), a distance vector is established that takes into account the longitudinal and lateral distances as well as the shape of the vehicle, and then the lateral and longitudinal safety distances are calculated according to the motion state of the vehicle, and the risk assessment function is obtained by combining them; The distance vector between the vehicle and the surrounding vehicles is expressed as: The longitudinal component r of the distance vector along the vehicle's velocity si , transverse component r li for: Where s represents the longitudinal position of the vehicle in the Frenet coordinate system, l represents the lateral position of the vehicle in the Frenet coordinate system, and θ i Indicates the heading angle of the vehicle numbered i, i represents the vehicle id, i=0,1,2,…,n, i=0 represents the ego vehicle, i=1,2,…,n represents the surrounding vehicles; s0 represents the longitudinal position of the ego vehicle in the Frenet coordinate system, l0 represents the lateral position of the ego vehicle in the Frenet coordinate system; s i represents the longitudinal position of the surrounding vehicles in the Frenet coordinate system, l i represents the lateral position of the surrounding vehicles in the Frenet coordinate system, i = 1, 2, ..., n; When considering the shape of the vehicle, the four corner points of the vehicle are used to represent the boundaries of the vehicle. According to the positional relationship between the surrounding vehicles and the vehicle itself, the longitudinal and lateral distances between the four corner points of the vehicle and the surrounding vehicles are obtained. The expressions are as follows: Where r represents the longitudinal and lateral distances between the four corners of the vehicle and the surrounding vehicles, a and b represent the length and width of the surrounding vehicles respectively; w and l represent the length and width of the vehicle respectively; poslin() function is a positive linear function; ± represents the positional relationship between the vehicle and the surrounding vehicles; S = {[1,1] T ,[-1,1] T ,[-1,-1] T ,[1,-1] T }; R is the rotation matrix, Indicates the yaw angle deviation of the vehicle and surrounding vehicles; Calculate the corresponding safety distances for the longitudinal and lateral directions respectively, and make an accurate assessment of the danger level of the vehicle through the longitudinal and lateral safety distances and the distance vector between vehicles; Longitudinal safety distance D si It is expressed as: In the formula, v0 represents the speed of the vehicle; v i represents the speed of surrounding vehicles; sign() is the sign function; ζ represents the reaction time, including the decision system reaction time and the brake action time; G represents the expected distance between the vehicle and surrounding vehicles; a min Indicates the minimum deceleration of the vehicle; θ i Indicates the heading angle of the vehicle; b1=v i ζ+G, lateral safety distance D li It is expressed as: Danger assessment function F i It is expressed as: The result of the hazard assessment function is affected by the speed and heading angle information of the vehicle and surrounding vehicles, forming a hazard field. The uncertainty of vehicle behavior is represented in the form of a hazard field, and the hazard area generated by the vehicle is corrected according to the maximum acceptable hazard as the final hazard domain.

6. The end-to-end decision-making planning method for the driving path of an autonomous driving vehicle according to claim 1, characterized in that: The step 3) specifically includes: Using all the candidate lanes obtained in step 1), after determining the candidate lanes according to the behavior of the vehicle, the future trajectories of all vehicles on the candidate lanes are predicted to obtain the position and speed information of all vehicles at several discrete moments in the future; through the hazard assessment function in step 2), the hazard assessment result at the corresponding moment is obtained, and combined with the boundary of each lane, a discrete time-space distance distribution map is obtained.

7. The end-to-end decision-making planning method for the driving path of an autonomous driving vehicle according to claim 1, characterized in that: In the step 4), an autonomous driving vehicle decision network is built based on a Transformer neural network, and a tree search strategy is combined to train and obtain the optimal target spacing; specifically, the following steps are included: 41) Encode the input data; Input data representation: Based on the Transformer neural network architecture, input data is received, including: vehicle history trajectory and environment information; for each vehicle, its history trajectory is a series of dynamic features in the past time step T = 20, including the vehicle's two-dimensional position, heading angle, speed and boundary size; for environment information, two types of map elements are considered, including lanes and crosswalks, each map element is represented by a vector, which consists of a series of nodes with different associated features; for each vehicle, a local environment information is constructed; the position attributes of all vehicles and map elements are converted to the local coordinate system of the AV; Lane line information encoding: The lane lines are discretized and sampled at a certain interval. Each sampling point is represented by the coordinates (x, y), and features are added, including line type and curvature. The features are normalized and spliced ​​into a multidimensional vector. Position encoding is used to add position information to each sampling point. Sine-cosine position encoding is used. The encoding formula is as follows: Among them, PE is the abbreviation of Positional Encoding, pos represents the position in the sequence, the dimension of the even-numbered bits uses 2i and the sin function, the dimension of the odd-numbered bits uses 2i+1 and the cos function, and d represents the total number of dimensions; Encode lane lines into a sequence to fit the input requirements of the Transformer neural network architecture; Vehicle history trajectory information encoding: The vehicle history trajectory is recorded in time series, including the vehicle position, speed, and acceleration; the data at different times are sequentially spliced ​​into vectors, and the time steps are t1, t2, ..., t n The corresponding positions are (x1,y1),(x2,y2),...,(x n ,y n ), speed v1,v2,...,v n Combine them into a long vector; add position encoding to make the vector have a time sequence and convert it into a position-aware sequence that can be processed by the Transformer neural network architecture; Environment Information Encoding: The environment information encoder consists of a lane encoder for processing lanes and a crosswalk encoder for processing crosswalk vectors; the lane encoder uses a multi-layer perceptron to encode numerical features and an embedding layer to encode discrete features; the crosswalk encoder is another multi-layer perceptron that encodes numerical features; 42) Select query, key, and value; Query Q selection: The vehicle’s position coordinates, speed, and acceleration information over the past multiple time steps are encoded as part of Q. Environmental information related to the vehicle’s current position, including the curvature of surrounding roads, the status of traffic signs and lights, and the relative positions of other vehicles, is embedded and represented as part of Q, so that the Transformer neural network can take into account the impact of environmental factors on the vehicle’s trajectory when making predictions. Key K selection: The detailed information in the vehicle's historical trajectory is taken as K, including the vehicle's steering angle changes and driving direction at different times; the environmental information is further encoded as K, including the topological structure information of the road, the location and type of surrounding buildings; Selection of value V: The value V of the vehicle's historical trajectory includes the average value of the vehicle's speed over a period of time, the standard deviation of the acceleration, and the short-term trend information predicted based on the vehicle's historical trajectory; The value V of the environmental information is the result of the fusion and abstraction of environmental features, including the aggregation of traffic flow conditions within a certain range of the surrounding area; The output of the attention mechanism is calculated by taking the query, key, and value as input, as follows: Among them, d k is the dimension of key K, K T represents the transpose of the K matrix; The results are passed through softmax to normalize the scores of all input quantities, so that the scores are all positive and sum to 1; 43) Coding process; Embedding layer: For the vehicle's historical trajectory, use the linear layer in the Transformer neural network to map it to a high-dimensional space to obtain a T×d matrix, where d is the embedding dimension; for environmental information, use a linear layer to map it to a d-dimensional vector; Multi-head attention layer: Q, K, and V generated by the encoding sequence are input into the multi-head attention module built into the multi-head attention layer of the Transformer neural network architecture; multiple heads calculate attention in parallel, and different heads focus on different subspace features of the input to capture the complex shape of lane lines, trajectory speed fluctuations, and circuitous navigation routes; Feedforward network layer: The output of the multi-head attention layer is input into the feedforward neural network built into the feedforward network layer. The feedforward neural network is a structure consisting of two fully connected layers and an activation function, which is used to perform nonlinear transformation on features, amplify important decision signals, suppress noise, and make lane lines, vehicle driving trajectories, and self-vehicle navigation information more decision-recognizable; the expression of the feedforward neural network is: FFN(x) = max(0,xW1+b1)W2+b2, where W1 and W2 are weight matrices, b1 and b2 are bias vectors, and FFN represents the feedforward neural network; 44) decoding process; Masked Multi-Head Self-Attention Layer: Use a masked multi-head self-attention layer to ensure that the neural network only pays attention to the previous time steps when generating predictions for each time step. Specifically, create a mask M where the elements on and below the diagonal are 0 and the elements above the diagonal are negative infinity. The mask matrix M is a 20×20 lower triangular matrix where the elements on and below the diagonal are 0 and the elements above the diagonal are negative infinity. The mask matrix M is added to the attention score matrix that has not been processed by softmax. When calculating softmax, the elements above the diagonal will approach 0, thus ensuring that the model will not be affected by future information when generating the output of the current time step; Cross attention layer: Attention calculation is performed between the output of the decoder and the output of the encoder to fuse the information of the encoded part. The required Q matrix comes from the output of the decoder, and the K matrix and V matrix come from the output of the encoder. Feedforward Neural Network: The output of the encoder-decoder attention layer passes through a feedforward neural network to obtain the final output of the decoder.

8. The end-to-end decision-making planning method for the driving path of an autonomous driving vehicle according to claim 7, characterized in that: The optimal target distance in step 4) is calculated by comprehensively considering the cost size, which includes calculating the arrival cost and efficiency cost of each distance, where: Arrival cost C arrival Including overlap cost C overlap and expansion cost C extend ; Overlap cost C overlap It is expressed as: in, Δv=v s -v e Among them, τ is the time corresponding to a lane change behavior, v s represents the speed of surrounding vehicles, v e Indicates the speed of the vehicle; Extension cost C extend The expression is: in, Indicates the front end position of the distance between the vehicle and the vehicle. Indicates the rear end position of the vehicle. Indicates the length of the gap between the vehicles. Indicates the front end position of the target spacing, Indicates the rear end position of the target spacing; Arrival cost C arrivall The expression is: C arrivall =C avarlap +C extend ; Efficiency cost C eff The expression is: Where V limit Indicates the speed limit of the lane where the gap is located, V front Indicates the speed of the vehicle ahead, V rear Indicates the speed of the vehicle behind.

9. The end-to-end decision-making planning method for the driving path of an autonomous driving vehicle according to claim 8, characterized in that: In the step 4), the optimal target distance adopts a tree search strategy to determine the transfer relationship between distances, specifically including taking the current distance of the vehicle as the root node of the tree, and the distance at each moment corresponds to the node of each layer on the search tree. If there is a transfer relationship between two distances, the corresponding two nodes are connected. At this time, the distance of another node other than the root node is taken as the optimal target distance; the Transformer neural network is verified according to the target distance sequence using the existing human driving data set, and then outputs the optimal target distance.

10. The end-to-end decision-making planning method for the driving path of an autonomous driving vehicle according to claim 9, characterized in that: The process of solving the driving trajectory in step 5) includes: According to the state of the vehicle in front of the target distance, the corresponding longitudinal point set is generated based on the intelligent driver model, and then the corresponding transverse point set is generated according to the transverse width of the space-time safety corridor. The longitudinal point set and the transverse point set are combined into a point set. The number of points in the point set is p = 1, 2, 3...n. Each point in the point set has a longitudinal coordinate s p and a transverse coordinate l p ; In the dynamic planning process, three costs are considered to ensure safety and comfort, including the smoothing cost C smooth , confined to the corridor, cost C bounds and the centering cost C middle ; Cost function C in the dynamic programming process total for: C smooth =abs((l p -l p-1 ) / (s p -s p-1 )) C middle =(l max +l min ) / 2 Among them, p represents a point in the point set, s p ,s p-1 represents the vertical coordinates of the pth point and the p-1th point, l p ,l p-1 represents the horizontal coordinates of the pth point and the p-1th point, l max , l min They represent the maximum and minimum values ​​of the corresponding time boundaries respectively; inf represents infinity, abs represents solving the absolute value; ω smooth ,ω bounds ,ω middle is the weight of the corresponding item; The reference trajectory is optimized using a quadratic programming algorithm. In the process of establishing the optimization equation, i.e., the objective function, the minimize jerk method is used; the optimization variables are selected. Among them, χ represents the optimization variable, s represents the longitudinal coordinate of the sampling point, and l represents the transverse coordinate of the sampling point. They represent the first-order, second-order, and third-order derivatives of the vertical coordinate with respect to time, respectively. They represent the first-order derivative, second-order derivative, and third-order derivative of the horizontal coordinate with respect to time respectively; The objective function consists of the closeness to the reference trajectory and comfort, where the closeness to the reference trajectory is expressed as The comfort cost includes the speed cost Acceleration cost and jerk cost The complete objective function is expressed as: Where g represents the reference trajectory; p represents the reference point; n represents the total number of sampling points; χ i ,x i ',x i ”, χ i ”'represent the pth optimization variable and the first-order, second-order, and third-order derivatives of the pth optimization variable, respectively. χ ,ω χ' ,ω χ” ,ω χ”' Respectively represent the weights of proximity, speed cost, acceleration cost, and comfort cost; The optimization equation is organized into a quadratic programming problem and solved using the QSQP solver to obtain the final driving trajectory.

Citation Information

Cited By

  • Vehicle automatic driving control method

    CN120817097A

  • Automatic driving lane changing trajectory planning method based on deep learning

    CN120963695A

  • Hybrid decision arbitration method and system based on automatic driving

    CN121246855A

  • Multi-target dynamic weight path planning method based on confidence feedback

    CN121632196A