An end-to-end autonomous driving decision planning model and device combining vectorized map and multi-task feature embedding
By combining multi-task feature embedding and cross-fusion of vectorized maps and bird's-eye view object detection, the robustness and interpretability issues of traditional autonomous driving models in dynamic environments are solved, and efficient integrated planning of perception and decision-making is achieved.
Patent Information
- Application Number
- CN202411078274.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-08-07
AI Technical Summary
Traditional autonomous driving decision-making and planning algorithms lack robustness and safety in highly dynamic and uncertain environments. Furthermore, existing end-to-end models are difficult to converge and lack interpretability, and error propagation and conflicts exist between tasks in multi-task models.
We adopt a network paradigm of imitation learning, combining targets in vectorized maps and bird's-eye view spaces as perception-aided tasks. Through multi-task feature embedding and multi-temporal cross-fusion modules, we construct an end-to-end autonomous driving decision-making and planning model. We use ResNet-50 feature extraction and cross-attention mechanism to fuse perception and decision information, and perform decision-making and planning through a Transformer decoder.
It improves the model's environmental awareness and the safety and stability of decision-making and planning, reduces error propagation between tasks, and enhances the model's interpretability and prediction accuracy.
Smart Images

Figure CN119116999B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of intelligent vehicle automatic driving, and relates to an automatic driving decision planning model and equipment based on vectorized maps and multi-task feature embedding. BACKGROUND
[0002] With the further development of artificial intelligence and intelligent transportation, automatic driving has become a research hotspot. The decision and planning module, as a core component of the automatic driving system, is a direct manifestation of the intelligence of automatic driving, and is responsible for formulating future behavior strategies and trajectory planning in combination with perception information and forward direction. The traditional rule-based automatic driving decision planning algorithm is difficult to guarantee the robustness and decision safety of the algorithm when facing high dynamicity, strong interaction and uncertainty in open road scenes. Therefore, a series of decision planning algorithms based on deep learning neural networks are widely used in the field of intelligent driving control.
[0003] The current common automatic driving system usually optimizes perception and decision planning as separate tasks, which adopts a layer-by-layer progressive architecture coupling, and the logic between different tasks is clear and the interpretability is strong. However, such methods also have certain deficiencies, for example, separate optimization between different modules, which leads to error transmission between modules and makes it difficult to achieve global optimization; at the same time, the perception task and the decision module exist separately, which leads to conflicts between the perception task output and the decision module, further leading to the disconnection of perception and decision. To overcome the above problems, a series of scientific researchers propose an end-to-end automatic driving decision planning, which directly uses sensor information as model input and outputs future behavior strategies and planning trajectories, and uses a deep learning model to optimize the perception and decision planning modules at the same time. However, simple end-to-end automatic driving models have the shortcomings of difficulty in model convergence and lack of model interpretability. To further overcome the above problems, relevant scholars propose an end-to-end automatic driving decision model based on multi-task learning, which introduces perception and decision-related auxiliary tasks and adds perception-related constraints to improve the model's ability to perceive the surrounding environment while increasing the model's output interpretability. For such multi-task automatic driving networks, how to select appropriate auxiliary tasks to perceive environmental information and how to construct an efficient multi-task coupling module are the key and difficult points in constructing a safe and efficient decision planning model.
[0004] In view of this, the application proposes an end-to-end automatic driving decision planning model combining vectorized maps and multi-task feature embedding. A network paradigm of imitation learning is adopted, the target in the vectorized map and bird's eye view space is combined as a perception auxiliary task, a multi-task feature embedding method is proposed to couple the perception and decision models, so as to realize a decision planning model that fully considers the surrounding environmental factors, and a multi-time sequence cross-fusion module is further proposed to use historical multi-frame information to further enhance the safety and stability of the decision model. SUMMARY
[0005] The application proposes an automatic driving decision planning model based on vectorized map and multi-task feature embedding. It mainly includes the design of three parts: environment perception module, multi-time cross attention fusion module and decision planning prediction module. Specifically, it includes the following contents:
[0006] Step one, make imitation learning dataset: with CARLA-0.9.10. simulation simulator, set different traffic scenes and driving tasks, and use expert driving algorithm combined with high-precision map for data collection. The main sensor information includes six ring-view cameras and GPS navigation. The collected data types mainly include: image information collected by six different cameras (I0, I1, I2, I3, I4, I5), each image size is (256, 256, 3); current speed information; next time target point; future T time step navigation point information obtained by expert experience and the expected speed output corresponding to each time; perception task label: including high-precision map vector information within the lateral distance [32m, -32m] and longitudinal distance [32m, -32m] at each time and other traffic participant position information in the bird's eye view space within this interval, semantic segmentation corresponding label of front view and depth map corresponding label.
[0007] Step two, build the environment perception module, which mainly includes three parts: 1) build the feature extraction module; 2) build the depth-semantic-image cross attention fusion module; 3) build the vectorized map and bird's eye view target detection task branch.
[0008] Step three, for the feature extraction module in step two, the application uses ResNet-50 as the main feature extraction network, and fuses multi-scale image features through the pyramid module. When the input image is (I0, I1, I2, I3, I4, I5), the corresponding feature maps extracted by the feature extraction module are (f0, f1, f2, f3, f4, f5).
[0009] Step four, for the depth-semantic-image cross attention fusion module in step two, the feature maps (f0, f1, f2, f3, f4, f5) collected in step three are respectively input into the depth information extraction network and the semantic information extraction network to obtain the corresponding depth information feature and semantic information feature To ensure that the extracted information contains depth information and semantic information, a 1x1 dimensional convolution is used for dimension conversion, and the collected semantic segmentation label and depth information label are used for supervision, and then a self-attention mechanism is used to respectively depth information feature and semantic information feature Further feature encoding is performed, and the feature encoding processes of the three different information are similar. Taking image feature encoding as an example, the calculation formula is as follows:
[0010] Q = w q f i ,K = w k f i ,V = w v f i
[0011]
[0012] wherein w q , w k , w v represent weight parameters that need to be learned by the model respectively, Softmax represents an activation function, K T is a transpose matrix of K. Q represents query features of input f i , K represents corresponding key features of input f i , V represents corresponding value features of input f i , Q I represents output features of self-attention, and d represents feature dimension. For depth encoding and semantic encoding, the calculation process is similar. According to the above process, image encoding features Q I , depth encoding features Q d and semantic encoding features Q e may be obtained respectively.
[0013] In order to further fuse the above different information to realize the bird's eye view feature output of the explicit image feature, the depth feature and the semantic feature, the application first initializes the query vector Q s =[Q m ,Q o ] according to the vectorized map output and the target detection output, wherein Q m represents a vectorized map query vector, and the application sets its dimension as 100*256, which represents that there are at most 100 pieces of vectorized map information within a range of 32 m around the assumption that needs to be defined, Q o represents a bird's eye view target detection query vector, and the application sets its dimension as 50*256, which represents that there are at most 50 other traffic participants within a range of 32 m around the assumption. First, the initialized semantic encoding feature query vector Q s is subjected to a self-attention mechanism, and the calculation process is as follows:
[0014] Q0 = w q0 Q s ,K0 = w k0Q s V0 = w v0 Q s
[0015]
[0016] wherein w q0 , w k0 , w v0 represent the weight parameters that the model needs to learn, Q0 represents the query feature corresponding to Q s , K represents the key feature corresponding to Q s , V represents the value feature corresponding to the input Q s , Softmax represents an activation function, and Q1 is the output query feature. Subsequently, the Q1 and the deep feature Q d are cross-attention fused:
[0017] Q'1 = w q1 Q1, K'1 = w k1 Q d , V'1 = w v1 Q d
[0018]
[0019] wherein w q1 , w k1 , w v1 represent the weight parameters that the model needs to learn, Q'1 represents the query feature corresponding to Q1, K'1 represents the key feature corresponding to Q d , V'1 represents the value feature corresponding to the input Q d , and Q2 represents the output query feature. The output query feature Q2 obtained above and the semantic feature Q e are cross-attention fused:
[0020] Q'2 = w q2 Q2, K'2 = w k2 Q e , V'2 = w v2 Q e
[0021]
[0022] wherein w q2 , w k2 , w v2 represent the weight parameters that the model needs to learn, Q'2 represents the query feature corresponding to Q2, K'2 represents the key feature corresponding to Q e , V'2 represents the value feature corresponding to Q eThe corresponding value feature is Q3, and Q3 represents the output query feature. The output query feature Q3 obtained above is fused with the image feature Q I Cross-attention fusion is performed:
[0023] Q'3 = w q3 Q3, K'3 = w k3 Q I V'3 = w v3 Q I
[0024]
[0025] wherein w q3 , w k3 , w v3 represent the weight parameters that the model needs to learn, Q'3 represents the corresponding query feature of Q3, K'3 represents the corresponding key feature of Q s , and V'3 represents the corresponding value feature of Q I Q4 represents the output query feature. Through the cross-attention mechanism described above, the query feature output Q4 that is fully fused with depth information and semantic information can be obtained. Then, Q4 is propagated through the forward propagation neural network to increase the output nonlinearity of the model, so that the final query feature output Q F is obtained.
[0026] Step five, for the vectorized map and the bird's eye view target detection task branch in step two, construct the output head of different tasks, and use the query feature Q F as input. Q F is divided into Q o and Q m according to the definition in step four, wherein the target detection output head uses a fully connected neural network combined with the query feature Q o to predict other surrounding traffic participants, and the output is wherein N1 = 50 represents the maximum number of other surrounding traffic participants, and the output P o contains the predicted category of the traffic participant, the length and width in the bird's eye view space. For the vectorized map, a fully connected neural network combined with the query feature Q m is used to predict the surrounding vector map information, and the output is wherein N2 = 100 represents the maximum prediction number of the surrounding vector map information, and the output P m contains the category information of the vectorized map and the prediction position information of ten points.
[0027] Step six, build a multi-time cross-attention fusion module. After obtaining the query feature Q FTo further ensure its effectiveness, the application cross-attention fuses it with the feature information of the historical moment to realize the feature output of the explicit historical information. Specifically, the Q F cross-attention feature fusion with the historical moment feature Q h is performed, and the calculation method is as follows:
[0028] Q′ F = w qf Q F ,K′ h = w kh Q h ,V′ h = w vh Q h
[0029]
[0030] where w qf , w kh , w vh represent the weight parameters that the model needs to learn, Q′ F represents the query feature corresponding to Q F , K′ h represents the key feature corresponding to Q h , V′ h represents the value feature corresponding to the input Q h , and Q″ F is the fused output feature. According to the above cross-attention mechanism, the fusion of the current moment feature and the historical moment feature can be obtained, and then the above feature Q″ F is input into the forward feedback neural network to obtain the final agent feature output Q a , and the historical moment feature Q h is updated as Q a .
[0031] Step seven, build a decision planning and prediction module, the input is the agent feature obtained after time series fusion in step six, mainly including two parts: 1) expected speed prediction model, 2) navigation point prediction model.
[0032] Step eight, for the expected speed prediction model in step seven, a model decoder based on Transformer is used. Specifically, first, to obtain the feature used to predict the expected speed, set the expected speed query feature and perform random initialization, and cross-fuse it with the obtained agent feature Q a , and the calculation method is as follows:
[0033] Q′ vp = w qa Qvp K' vp = w ka Q a V' vp = w va Q a
[0034]
[0035] where w qa , w ka , w va represent the weight parameters that the model needs to learn, Q' vp represents the Q vp corresponding query feature, K' vp represents the Q a corresponding key feature, V' vp represents the input Q a corresponding value feature, Q" vp represents the output expected speed feature, and then Q" vp is output to the front feedback neural network and the fully connected neural network to obtain the final expected speed prediction.
[0036] Step nine, for the navigation point prediction model in step seven, a Transformer-based model decoder is also used. Specifically, first, to obtain the features used to predict the navigation point, set the navigation point query feature and randomly initialize it, where T represents the future prediction time step of the navigation point, which is set to 10 in the present application. Then input it into the self-attention module, which is calculated as follows:
[0037] Q' wp = w qw Q wp K' kw = w ka Q wp V' vw = w va Q wp
[0038]
[0039] where w qw , w kw , w vw represent the weight parameters that the model needs to learn, Q' wp , K' kw , V' vw represent the input Q wp corresponding query feature, key feature and value feature, respectively, Q" wp is the output query feature, and it is combined with the obtained agent feature Qa Cross fusion is performed, which is calculated as follows:
[0040] Q wp = w qw Q wp K kw = w kw Q a V vw = w vw Q a
[0041]
[0042] wherein w qw , w kw , w v represent the weight parameters to be learned by the model, Q wp represents Q wp corresponding query features, K kw represents Q a corresponding key features, V vp represents input Q a corresponding value features, represent the output navigation point features, and then output to the forward feedback neural network to increase the nonlinearity of the features and input to the GRU time series prediction model to obtain the final navigation point prediction.
[0043] Step ten, for the GRU time series prediction model in step nine: composed of ten GRU prediction units, each GRU prediction unit combines a fully connected neural network to output a navigation point output corresponding to a time step. Specifically, for the navigation point features input, it is divided into navigation point features of different time steps tar Input the target point w to the fully connected neural network to obtain 1 as the initial hidden feature. Then input the initial hidden feature and the initial time step navigation point feature to the first GRU prediction unit to obtain the output feature h 0 , which is used as the hidden feature at the next time, and the output feature h 0 is input to the fully connected neural network to obtain the navigation point prediction at the initial time , and so on, so as to obtain the navigation point outputs of the next ten time steps
[0044] Step eleven, the control execution module is built, the expected speed prediction obtained in step eight and the navigation point prediction obtained in step nine are taken as inputs, a control model is constructed, and the control model mainly includes two parts, i.e., a longitudinal control part for controlling the throttle and the brake and a lateral control part for controlling the steering wheel angle.
[0045] Step twelve, for longitudinal control, a threshold T is first set v When the predicted expected speed is less than T v , the brake at this time is set to 1 to realize brake control, otherwise the predicted expected speed is input into a PID controller to obtain the throttle opening at this time to realize acceleration control.
[0046] Step thirteen, for lateral control, the average value of the predicted output of the navigation point is first calculated, and the average value is taken as the inverse tangent to obtain the predicted angle A p at this time.
[0047]
[0048] Where (x i , y i ) are the coordinates of the navigation point , and the predicted angle A p obtained is input into a PID controller to obtain the steering wheel angle u p .
[0049] Step thirteen, the end-to-end automatic driving multi-task decision planning model constructed above is trained. Different loss functions are constructed for different task modules, wherein the depth estimation task, the expected speed prediction and the semantic segmentation adopt the L1 loss function, the semantic segmentation adopts the cross-entropy loss function, and for the target detection and the vectorized map construction, the loss function includes two parts: the cross-entropy loss function of the category prediction, and the L1 loss function of the regression part. Thus, the overall loss function L is:
[0050]
[0051] Wherein, respectively represent the semantic segmentation, the depth estimation, the vectorized map construction, the target detection, the navigation point prediction and the expected speed prediction loss function.
[0052] The beneficial effects of the present application are:
[0053] (1) The present application proposes an end-to-end automatic driving decision planning model based on multi-task coupling, and six task branches including perception, decision and control are designed. The end-to-end automatic driving framework is used to reduce the error between different model transmissions, and the multi-task constraint method is used to increase the fitting ability of the model.
[0054] (2) The application combines a vectorized map with bird's eye view target detection as a perception representation of a multi-task model. Dynamic and static target information required in a bird's eye view space is provided, and environment perception information and decision-making navigation point output are unified to the same space, thereby improving model prediction accuracy.
[0055] (3) The application adopts a multi-task feature embedding coupling mode, uses a cross attention mechanism between different tasks to transmit information, overcomes training conflicts caused by simple shared feature maps in traditional multi-task models, and realizes efficient coupling between different tasks. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 is a vectorized map combined with a multi-task feature embedding network architecture diagram;
[0057] Figure 2 is a control model structure diagram; DETAILED DESCRIPTION
[0058] The application will be further described below in conjunction with the accompanying drawings and specific embodiments, but the protection scope of the application is not limited thereto.
[0059] Figure 1 is a vectorized map combined with a multi-task feature embedding network architecture diagram, mainly including three parts: a perception module, a multi-time processing module, and a decision output model, which are specifically as follows:
[0060] (1) Making imitation learning dataset: Making imitation learning dataset: With the help of CARLA-0.9.10. simulation simulator, different traffic scenes and driving tasks are set, and a rule-based expert driving algorithm combined with a high-precision map is used for data collection. The main sensor information includes six ring-view cameras and GPS navigation. The collected data types mainly include: image information (I0, I1, I2, I3, I4, I5) collected by six different cameras, each image size is (256, 256, 3); current speed information; next time target point; future T time step navigation point information obtained by expert experience and the expected speed output corresponding to each time; perception task label: including high-precision map vector information within the lateral distance [32m, -32m] and longitudinal distance [32m, -32m] at each time and other traffic participant position information in the bird's eye view space within the interval, semantic segmentation corresponding label of the front view and label corresponding to the depth map.
[0061] (2) Build an environment perception module, mainly including three parts: 1) build a feature extraction module; 2) build a deep-semantic-image cross attention fusion module; 3) build a vectorized map and bird's eye view target detection task branch.
[0062] (3) The application adopts ResNet-50 as the main feature extraction network, and fuses multi-scale image features through a pyramid module, when the input image is (I0, I1, I2, I3, I4, I5), the corresponding feature maps extracted through the feature module are (f0, f1, f2, f3, f4, f5).
[0063] (4) Build a deep-semantic-image cross attention fusion module, input the collected feature maps (f0, f1, f2, f3, f4, f5) into a depth information extraction network and a semantic information extraction network respectively, and obtain corresponding depth information features and semantic information features To ensure that the extracted information contains depth information and semantic information, 1*1 dimensional convolution is used for dimension conversion, and the collected semantic segmentation labels and depth information labels are combined for supervision, and then a self-attention mechanism is used to further encode the image features depth information features and semantic information features The feature encoding processes of the three different information are similar, and the calculation formula of the image feature encoding is as follows:
[0064] Q=w q f i ,K=w k f i ,V=w v f i
[0065]
[0066] Wherein, w q , w k , w v respectively represent the weight parameters that the model needs to learn, Softmax represents the activation function, K T is the transpose matrix of K. Q represents the query feature of input f i , K represents the corresponding key feature of input f i , V represents the corresponding value feature of input f i , Q I represents the output feature of self-attention, d represents the feature dimension. For depth encoding and semantic encoding, the calculation process is similar, and according to the above process, the image encoding feature Q I , the depth encoding feature Qd and semantic encoding feature Q e .
[0067] To further fuse the above different information to realize the bird's eye view feature output of explicit image feature, depth feature and semantic feature, the application first initializes the query vector Q according to the vectorized map output and the target detection output s =[Q m ,Q o ], wherein Q m represents the vectorized map query vector, and the application sets its dimension as 100*256, representing that there are at most 100 pieces of vectorized map information to be defined within the surrounding range of 32 m, and Q o represents the target detection query vector under the bird's eye view, and the application sets its dimension as 50*256, representing that there are at most 50 other traffic participants within the surrounding range of 32 m. First, the initialized query vector Q s passes through the self-attention mechanism, and the calculation process is as follows:
[0068] Q0=w q0 Q s ,K0=w k0 Q s ,V0=w v0 Q s
[0069]
[0070] wherein w q0 , w k0 , w v0 represent the weight parameters to be learned by the model, Softmax represents the activation function, and d represents the dimension number of the feature. Then, the output query feature Q1 is cross-attention fused with the depth feature Q d .
[0071] Q′1=w q1 Q1,K′1=w k1 Q d ,V′1=w v1 Q d
[0072]
[0073] wherein w q0 , w k0 , w v0 represent the weight parameters to be learned by the model, Q0 represents the corresponding query feature of Q s , K represents the corresponding key feature of Q s , and V represents the input Q sCorresponding value feature, Softmax represents the activation function, d represents the dimension number of the feature, Q1 is the output query feature. Then Q1 is fused with the deep feature Q d Cross attention fusion is performed:
[0074] Q'1 = w q1 Q1, K'1 = w k1 Q d V'1 = w v1 Q d
[0075]
[0076] wherein, w q1 , w k1 , w v1 represent the weight parameters that the model needs to learn, Q'1 represents the corresponding query feature of Q1, K'1 represents the corresponding key feature of Q d , and V'1 represents the corresponding value feature of the input Q d Q2 represents the output query feature. The output query feature Q2 obtained above is fused with the semantic feature Q e Cross attention fusion is performed:
[0077] Q'2 = w q2 Q2, K'2 = w k2 Q e V'2 = w v2 Q e
[0078]
[0079] wherein, w q2 , w k2 , w v2 represent the weight parameters that the model needs to learn, Q'2 represents the corresponding query feature of Q2, K'2 represents the corresponding key feature of Q e , and V'2 represents the corresponding value feature of the input Q e Q3 represents the output query feature. The output query feature Q3 obtained above is fused with the image feature Q I Cross attention fusion is performed:
[0080] Q'3 = w q3 Q3, K'3 = w k3 Q I V'3 = w v3 Q I
[0081]
[0082] wherein, w q3 , wk3 , w v3 represent the weight parameters that the model needs to learn, Q'3 represents the query feature corresponding to Q3, K'3 represents the key feature corresponding to Q s , and V'3 represents the value feature corresponding to Q I . Through the above cross-attention mechanism, the query feature output Q4 fully fused with depth information and semantic information can be obtained, and then Q4 is propagated through the forward propagation neural network to increase the output nonlinearity of the model, so that the final query feature output Q F is obtained.
[0083] (5) Construct a vector map and an aerial view target detection task branch, adopt the query feature Q F as input, divide Q F into Q o and Q m according to the definition in step four, wherein the target detection output head adopts a fully connected neural network combined with the query feature Q o to predict other surrounding traffic participants, and the output is wherein N1=50 represents the maximum number of other surrounding traffic participants, and the output contains the category, the length and the width in the aerial view space. For the vector map, a fully connected neural network combined with the query feature Q m is adopted to predict the surrounding vector map information, and the output is wherein N2=100 represents the maximum prediction number of the surrounding vector map information, and the output contains the category information of the vector map and the prediction position information of ten points.
[0084] (6) Build a multi-time cross-attention fusion module. After obtaining the query feature Q F that can effectively perceive the surrounding environment information, in order to further ensure its effectiveness, the application cross-attends and fuses the feature information at the historical moment to realize the feature output of the explicit historical information. Specifically, the cross-attention feature fusion is performed between the current moment Q F and the historical moment feature Q h , and the calculation method is as follows:
[0085] Q' F = w qf Q F , K' h = w kh Q h , V' h = w vh Q h
[0086]
[0087] wherein w qf , w kh , w vh represent weight parameters that need to be learned by the model, Q' F represents the query feature corresponding to Q F , K' h represents the key feature corresponding to Q h , V' h represents the value feature corresponding to Q h , and Q" F is the fused output feature. According to the cross-attention mechanism described above, the fusion of the current time feature and the historical time feature can be obtained, and then the above feature Q" F is input into the forward feedback neural network to obtain the final agent feature output Q a , and the historical time feature Q h is updated as Q a .
[0088] (5) A decision planning and prediction module is built, and the input is the agent feature obtained after time sequence fusion in step six, which mainly includes two parts: 1) an expected speed prediction model, and 2) a navigation point prediction model.
[0089] (7) A model decoder based on Transformer is used to build the expected speed prediction model. Specifically, first, the expected speed query feature Q is set to obtain the feature used to predict the expected speed, and is randomly initialized, and is cross-fused with the obtained agent feature, and the calculation method is as follows:
[0090] Q' vp = w qa Q vp , K' vp = w ka Q a , V' vp = w va Q a
[0091]
[0092] wherein w qa , w ka , w va represent weight parameters that need to be learned by the model, Q' vp represents the query feature corresponding to Q vp , K' vp represents the key feature corresponding to Q a , V' vp represents the value feature corresponding to Q a , and Q" vpThe expected speed feature of the output is represented, and then Q" vp is output to the front feedback neural network and the fully connected neural network to obtain the final expected speed prediction.
[0093] (8) A navigation point prediction model is constructed by using a Transformer-based model decoder. Specifically, first, to obtain the feature for predicting the navigation point, the navigation point query feature is randomly initialized, where T represents the future prediction time step of the navigation point, and the present application sets T to 10. Then, it is input into the self-attention module, which is calculated as follows:
[0094] Q' wp = w qw Q wp ,K' kw = w ka Q wp ,V' vw = w va Q wp
[0095]
[0096] where w qw , w kw , w vw represent the weight parameters that the model needs to learn, Q' wp , K' kw , V' vw represent the input Q wp corresponding query feature, key feature and value feature, Q" wp is the output query feature, and it is cross-fused with the obtained agent feature Q a , which is calculated as follows:
[0097] Q''' wp = w' qw Q" wp ,K" kw = w' kw Q a ,V" vw = w' vw Q a
[0098]
[0099] where w' qw , w' kw , w' vw represent the weight parameters that the model needs to learn, Q''' wp represents the Q" wp corresponding query feature, K"kw representing Q a corresponding key features, V' vp representing input Q a corresponding value features, representing output navigation point features, and then output to the forward feedback neural network to increase the nonlinearity of the features and input to the GRU time series prediction model to obtain the final navigation point prediction. The GRU time series prediction model is composed of ten GRU prediction units, and each GRU prediction unit combines a layer of fully connected neural network to output a navigation point output corresponding to a time step. Specifically, for its input navigation point features it is divided into navigation point features at different time steps adopt target point w tar input to the fully connected neural network to obtain as the initial hidden feature. Then the initial hidden feature and the initial time step navigation point feature are input to the first GRU prediction unit to obtain the output feature h 1 , which is taken as the hidden feature at the next time, and the output feature h 0 is input to the fully connected neural network to obtain the navigation point prediction at the initial time and so on, so as to obtain the navigation point output at the next ten time steps
[0100] (10) Build a control execution model, take the obtained expected speed prediction and navigation point prediction as input to construct a control model, mainly including two parts for controlling the throttle and brake longitudinal control and for controlling the steering wheel angle of the lateral control, the network structure is as shown in Figure 2 For longitudinal control, first set a threshold T v , when the predicted expected speed is less than T v , set the brake at this time to 1, otherwise input the predicted expected speed into the PID controller to obtain the throttle opening at this time. For lateral control, first take the average of the predicted output of the navigation point, and take the tangent to obtain the predicted angle A p at this time:
[0101]
[0102] where (x i , y i ) is the coordinate of the navigation point , and the predicted angle A p obtained is input to the PID controller to obtain the steering wheel angle u p .
[0103] (11) The end-to-end automatic driving multi-task model constructed above is trained. Different loss functions are constructed for different task modules, wherein the depth estimation task, the expected speed prediction and the L1 loss function are adopted, the semantic segmentation adopts the cross-entropy loss function, for the target detection and the vectorized map construction, the loss function includes two parts: the cross-entropy loss function of the category prediction, and the L1 loss function of the regression part. Thus, the overall loss function can be obtained is:
[0104]
[0105] wherein, respectively represent the semantic segmentation, the depth estimation, the vectorized map construction, the target detection, the navigation point prediction and the expected speed prediction loss function.
[0106] (12) The embodiment of the present application uniformly proposes an intelligent driving vehicle device, which utilizes the trained model to combine a vehicle-mounted industrial personal computer, a central controller and a camera and the like to realize the automatic driving decision planning described in the application.
[0107] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation modes of the present application, and are not used to limit the protection scope of the present application, and any equivalent mode or change without departing from the technology of the present application should be included in the protection scope of the present application.
Claims
1. An end-to-end autonomous driving decision planning model that combines vectorized maps with multi-task feature embeddings, characterized in that, Obtained by the following: S1, build an environment perception module, mainly including: feature extraction module, deep-semantic-image cross attention fusion module and vectorized map and bird's eye view target detection task branch module; The feature extraction module is used for extracting the features of the input image; The depth-semantic-image cross-attention fusion module obtains corresponding depth information features and semantic information features according to the features extracted by the feature extraction module, encodes the extracted image features, depth information features and semantic information features to obtain image encoding features, depth encoding features and semantic encoding features, and performs information fusion to obtain the final query feature Q F ; The vectorized map and bird's eye view target detection task branch module takes query features as input, constructs output heads for different tasks, and finally outputs class information, long and wide under the bird's eye view space, and ten point prediction position information of the vectorized map; S2, build a multi-time sequence cross attention fusion module to obtain query features Q that can effectively perceive surrounding environment information F After that, to further ensure its effectiveness, cross attention fusion is performed with the feature information at the historical moment to obtain the time sequence fused agent feature Q a To achieve a feature output that clearly indicates historical information; The specific implementation of S2 includes: Q F Q h cross attention feature fusion, which is calculated as follows: Q' F = w qf Q F ,K' h = w kh Q h ,V' h = w vh Q h wherein w qf , w kh , w vh represent the weight parameters to be learned by the model, Q' F represents the Q F corresponding query feature, K' h represents the Q h corresponding key feature, V' h represents the Q h corresponding value feature, Q" F is the fused output feature, according to the cross-attention mechanism described above, the fusion of the current time feature and the historical time feature can be obtained, and then the above Q" F is input into the forward feedback neural network to obtain the final agent feature output Q a , and the historical time feature Q h is updated as Q a ; S3, build a decision planning and prediction module, which mainly includes two parts: 1) expected speed prediction model, 2) navigation point prediction model; The expected speed prediction model cross-fuses the expected speed query feature and the time-series fused agent feature to obtain the expected speed prediction. The navigation point prediction model first performs self-attention operation on the navigation point query feature, then cross-fuses it with the agent feature, and performs GRU time-series prediction to obtain the navigation point prediction. 2.The end-to-end autonomous driving decision planning model combining vectorized map with multi-task feature embedding of claim 1, wherein, In S1: the feature extraction module uses ResNet-50 as the main feature extraction network, and fuses multi-scale image features through the pyramid module. When the input image is (I0, I1, I2, I3, I4, I5), the corresponding feature maps extracted by the feature extraction module are (f0, f1, f2, f3, f4, f5), respectively. A depth-semantic-image cross-attention fusion module, the feature maps (f0, f1, f2, f3, f4, f5) extracted by the feature extraction module are respectively input into a depth information extraction network and a semantic information extraction network to obtain corresponding depth information features and semantic information features To ensure that the extracted information contains depth information and semantic information, a 1x1 dimensional convolution is used for dimension conversion, and the collected semantic segmentation label and depth information label are combined for supervision, and then a self-attention mechanism is used to further encode the image features depth information features and semantic information features The feature encoding processes of the three different information are the same. To fuse the above different information to realize the bird's eye view feature output of clear image features, depth features and semantic features, first, initialize the query vector Q according to the vectorized map output and the target detection output s = [Q m , Q o ], wherein Q m represents the vectorized map query vector, and the dimension is set to 100*256, representing that there are at most 100 pieces of vectorized map information within a range of 32 m around the vehicle, and Q o represents the target detection query vector under the bird's eye view, and the dimension is set to 50*256, representing that there are at most 50 other traffic participants within a range of 32 m around the vehicle; first, the initialized query vector Q s is subjected to self-attention mechanism, and the calculation process is as follows: Q0 = w q0 Q s K0 = w k0 Q s V0 = w v0 Q s wherein w q0 , w k0 , W v0 represent weight parameters to be learned by the model, Q0represents the query feature corresponding to Q s , K0represents the key feature corresponding to Q s , V0represents the value feature corresponding to Q s , Softmax represents an activation function, and Q1is an output query feature. Subsequently, the output query feature Q1is cross-attention fused with the deep feature Q d : Q'1 = w q1 Q1,K'1 = w k1 Q d ,V'1 = w v1 Q d wherein w q1 , w k1 , w v1 represent weight parameters to be learned by the model, Q'1 represents the query feature corresponding to Q1, K'1 represents the key feature corresponding to Q d , V'1 represents the value feature corresponding to the input Q d , Q2 represents the output query feature, d represents the feature dimension, and the output query feature Q2 obtained above is cross-attention fused with the semantic feature Q e . Q'2= w q2 Q2, K'2= w k2 Q s ,V'2= w v2 Q s wherein w q2 , w k2 , w v2 represent weight parameters to be learned by the model, Q'2 represents the query feature corresponding to Q2, K'2 represents the key feature corresponding to Q e , and V'2 represents the value feature corresponding to Q e , Q3 represents the output query feature, and the output query feature Q3 obtained above is cross-attention fused with the image feature Q I . Q'3 = w q3 Q3, K'3 = w k3 Q I ,V'3 = w v3 Q I wherein w q3 , w k3 , w v3 represent weight parameters to be learned by the model, Q'3 represents the query feature corresponding to Q3, K'3 represents the key feature corresponding to Q s , V'3 represents the value feature corresponding to the input Q I , Q4 represents the output query feature, through the above cross-attention mechanism, the query feature output Q4 fully fused with the depth information and the semantic information can be obtained, and then Q4 is propagated through the forward propagation neural network to increase the output nonlinearity of the model, so that the final query feature output Q F is obtained. 3.The end-to-end autonomous driving decision planning model combining vectorized map with multi-task feature embedding of claim 2, wherein, The vectorized map in S1 and the bird's eye view target detection task branch module construct the output head of different tasks, and query the feature Q F As input, Q F is divided into Q o and Q m Among them, the target detection output head adopts a fully connected neural network combined with the query feature Q o Predict other surrounding traffic participants, and the output is Where N1=50 represents the maximum number of other surrounding traffic participants, and the output P o contains the category, the length and width under the bird's eye view space; for the vectorized map, a fully connected neural network combined with the query feature Q m Predict the surrounding vector map information, and the output is Where N2=100 represents the maximum prediction number of surrounding vector map information, and the output P m contains the category information of the vectorized map and the prediction position information of ten points. 4.The end-to-end autonomous driving decision planning model combining vectorized map with multi-task feature embedding of claim 1, wherein, The expected speed prediction model of the S3 adopts a Transformer-based model decoder, first sets an expected speed query feature for obtaining a feature for predicting the expected speed And randomly initialize it with the obtained agent feature Q a Cross fusion; The navigation point prediction model adopts a Transformer-based model decoder. First, a navigation point query feature is set to obtain a feature for predicting a navigation point in a prediction period and is randomly initialized, where T represents a future prediction time step of the navigation point, and then is input into a self-attention module. 5.The end-to-end autonomous driving decision planning model combining vectorized map with multi-task feature embedding of claim 1, wherein, It also includes S4, building a control execution module, which realizes the longitudinal control of the accelerator and brake and the lateral control of the steering wheel angle according to the expected speed prediction and navigation point prediction. For longitudinal control, first set a threshold T v When the predicted expected speed is less than T v , set the brake to 1 at this time to realize brake control; otherwise, input the predicted expected speed into the PID controller to obtain the corresponding throttle opening to realize acceleration control; For lateral control, first the average of the predicted outputs of the waypoints is taken, and the inverse tangent is taken to get the predicted angle A at that time p : Where T represents the future T time steps, (x i ,y i ) is the coordinate of the navigation point, and the predicted angle A is obtained again. The predicted angle A p is input into the PID controller to obtain the steering wheel angle u p . 6.The end-to-end autonomous driving decision planning model combining vectorized map with multi-task feature embedding of claim 1, wherein, It also includes making a dataset for model training: With the help of CARLA-0.9.
10. simulation simulator, different traffic scenes and driving tasks are set, and the data is collected by using the rule-based expert driving algorithm combined with high-precision maps. The set sensor information includes six look-around cameras and GPS navigators; the collected data types include: image information (I0, I1, I2, I3, I4, I5) collected by six different cameras, the size of each image is (256, 256, 3), the speed information of the current time, the target point of the next time, the navigation point information of the future T time steps obtained by expert experience and the expected speed output corresponding to each time, the perception task label: including the high-precision map vector information within the lateral distance [32m, -32m] and the longitudinal distance [32m, -32m] at each time and the position information of other road users in the bird's eye view space within this interval, the semantic segmentation corresponding label of the front view and the label corresponding to the depth map.
7. The end-to-end autonomous driving decision planning model incorporating vectorized map and multi-task feature embedding according to claim 6, characterized in that, Also comprising: training the model with the produced dataset, constructing corresponding loss functions for different task modules, wherein the depth estimation task and the expected speed prediction adopt an L1 loss function, the semantic segmentation adopts a cross-entropy loss function, and for the target detection and the vectorized map construction, the loss function comprises two parts: a cross-entropy loss function for class prediction and an L1 loss function for the regression part, so that the overall loss function can be obtained To: wherein, respectively represent semantic segmentation, depth estimation, vectorized map construction, object detection, navigation point prediction, and expected speed prediction loss functions.
8. An intelligent driving vehicle control device, characterized by comprising: The device deploys the trained end-to-end autonomous driving decision planning model as claimed in claim 7 combined with vectorized map and multi-task feature embedding.
Citation Information
Patent Citations
Metareinforcement learning-based end-to-end automatic driving method and system
CN116469080A
End-to-end automatic driving decision planning method and device in combination with meta-learning multi-task optimization
CN116729433A