Uncertain end-to-end automatic driving method, model and equipment combining distributed query and fusing spatio-temporal information
Through multi-scale hollow aggregation, distributed query and joint probability motion planning, the problems of high collision rate and large L2 error of deterministic end-to-end autonomous driving methods in complex scenarios are solved, and more efficient and safe autonomous driving is achieved.
Patent Information
- Application Number
- CN202510358654.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing deterministic end-to-end autonomous driving methods have problems of high collision rates and large L2 errors in dynamic and highly complex urban market scenarios. This is mainly because static scene representation and dynamic traffic participant representation follow the deterministic paradigm, which ignores the impact of static uncertainty and dynamic uncertainty during planning, and is inefficient in resource efficiency.
Multi-scale hollow aggregation, distributed query mechanism, scene understanding module and joint probability motion planning module are adopted to enhance feature representation through multi-scale hollow convolution, distributed query decouples task dependence, and Laplace distribution modeling static uncertainty, joint probability motion planning model dynamic uncertainty, and resource efficiency is improved through sparse representation.
It reduces the collision rate and L2 error, improves resource efficiency, enhances the interpretability and security of the planning, and can better handle dynamic and complex scenarios.
Smart Images

Figure CN120296559A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent vehicle autonomous driving, and relates to an end-to-end autonomous driving method, model and device that combines distributed query and fuses the uncertainty of spatio-temporal information. Technical Background
[0002] The rapid development of autonomous driving technology has promoted the progress of the transportation field. The autonomous driving technology combined with pure vision is widely used in the passenger vehicle assisted driving system due to its low cost and rich information, which is crucial for ensuring the safety of the autonomous driving system. Autonomous driving mainly includes tasks such as environmental perception, motion prediction, and trajectory planning. The traditional modular paradigm driven by rules independently trains and optimizes each module and then cascades and couples them. Although it has a certain degree of interpretability, the information loss and error accumulation during the coupling of modules severely limit the planning performance, resulting in safety problems.
[0003] The data-driven end-to-end paradigm provides an effective solution to solve the above problems. It regards perception, prediction, and planning as a whole and optimizes the entire system towards the final planning goal, significantly reducing the error accumulation during module coupling. Currently, it is mainly divided into the fully end-to-end paradigm and the modular end-to-end paradigm. The fully end-to-end paradigm directly obtains the planning trajectory or control signal using sensor input. This paradigm has no explicit perception and prediction for supervision, solves the problem of error accumulation, but cannot obtain the effects of safety, interpretability, and continuous optimization, especially for dynamic and highly complex urban scenarios. For example, directly extract the original BEV features and the ego-vehicle state using a query, and then use an MLP to obtain the final trajectory. The modular end-to-end paradigm uses an implicit query mechanism to connect all-stack auxiliary supervision tasks such as perception and prediction of autonomous driving into a network, and uses queries for interaction and optimization. This paradigm has stronger interpretability and scene understanding ability and is easy to optimize, but there is still a relatively high collision rate and L2 error during planning in dynamic and highly complex urban scenarios, seriously affecting safety.
[0004] The present invention deeply analyzes the above working structural principles and finds that both their static scene representations and dynamic traffic participant representations follow deterministic paradigms, and the feature extraction and task connection sequences also follow deterministic paradigms. This deterministic modeling paradigm will lead to serious collision rates and L2 errors during planning. In the feature decoding stage, when the end-to-end pipeline makes decision-making plans based on the dynamic trajectories of traffic participants and the rules of static map elements, it will ignore the impacts of static uncertainties and dynamic uncertainties in the driving scene on the planning task. For example, when performing trajectory planning based on a deterministic static scene, however, the mispositioning, offset, and misidentification of static scene map elements will cause the ego vehicle to drive following incorrect scene rules, resulting in deviations in the planned path; another example is assuming a deterministic relationship between the actions of all traffic participants and the environment and then performing trajectory planning. However, the driving styles of traffic participants themselves are highly variable, and it is unreasonable to plan only one path for highly variable trajectories. Modeling large uncertainties with this deterministic paradigm will inevitably result in high collision rates and L2 errors. Tracing back to the feature encoding stage, when the convolutional neural network performs feature extraction and aggregation, it will default that the importance of features in different channel dimensions and spatial dimensions is deterministically equivalent and assign equal weights. In fact, the contributions of different dimensions to the ultimate goal of autonomous driving are different, resulting in unreasonable information allocation, which will reduce the detection accuracy of obstacles and map elements and the accuracy of trajectory planning. Under this deterministic paradigm, the entire end-to-end pipeline sequentially executes environment perception, agent motion prediction, and ego vehicle planning in a deterministic order, resulting in the planning task only using the prior information provided by scene understanding and motion prediction, that is, the planning effect completely depends on upstream scene understanding and motion prediction. Once the prior information fails, it will inevitably lead to planning failure. Moreover, this deterministic order cannot model the impact of the ego vehicle's trajectory on agent motion, ignoring the high-order interaction effects of agents, maps, temporal information, and ego vehicle planning, which will greatly increase the impacts brought by external uncertainties. Additionally, based on the BEV dense representation method, the resource efficiency is low, accompanied by problems of high memory and computational costs, and it is difficult to expand the spatial range and utilize temporal information. Summary of the Invention
[0005] Based on the above background, the present invention proposes an uncertainty end-to-end autonomous driving method that combines distributed queries and fuses spatio-temporal information, as shown in the appendix Figure 1As shown in the figure, the following work has been carried out to achieve this goal. In the feature decoding stage, to solve the problem of deterministic static representation, the output of the scene decoder is modeled as a Laplace distribution, representing the uncertain static map elements in the form of a probability distribution and passing this uncertainty to the planning, so that the planning can fully consider the error brought by this uncertainty. To solve the problem of deterministic dynamic representation, the ego-vehicle planning actions are represented as a vocabulary group, and then the inherent uncertainty of the planning is modeled as a probability field distribution, and the trajectory with the highest probability is sampled from the multi-modal planning trajectories as the current output to fully reduce the impact of planning uncertainty. Further, to fully consider the impact of the ego-vehicle on other vehicles, using the similarity between prediction and planning, the agent motion prediction and the ego-vehicle planning are simultaneously modeled in the way of joint probability motion planning, and the agent, the map, the historical features and the ego-vehicle are interacted at a high level to fully consider the impact of the external environment uncertainty. In the feature encoding stage, to solve the problem of deterministic dimensional contribution, multi-scale atrous aggregation is proposed. The receptive field is enlarged by using multi-scale atrous convolution to enhance the local-global relationship, and the importance weights of features in different dimensions are dynamically adjusted by using channel aggregation and spatial aggregation to enhance the representation ability. For the end-to-end network, a task-specific distributed query is proposed to directly decode the original features, aiming to expand the source of planning information and relieve its dependence on the prior information of other multi-tasks. Taking planning as the main task and detection and mapping as auxiliary tasks, the auxiliary tasks supervise the learning of the main task to improve interpretability. A sparse representation method and a streaming time series strategy are adopted to improve resource efficiency.
[0006] The present invention proposes an uncertainty end-to-end autonomous driving method, model and device combining distributed query and fusing spatio-temporal information, and its main components include the following parts: 1. Feature encoding combining multi-scale atrous convolution. 2. Distributed query mechanism. 3. Scene understanding module. 4. Joint probability motion planning module.
[0007] The specific steps of the method are as follows:
[0008] Step 1, multi-scale atrous aggregation. First, use a convolutional neural network to extract image features, and then use multi-scale atrous aggregation to enhance the features to obtain the final scene features. Different dimensions of the convolutional neural network represent different features or patterns, and the importance of the information contained is different. However, when extracting features, it is default that the contribution importance of different dimensions is deterministically equal, resulting in unreasonable weight allocation. In addition, the dynamic traffic participants in the driving scene are very different in size and shape from the static map, and the fixed receptive field of the convolutional neural network limits its ability to capture features at different scales. To solve this problem, multi-scale atrous aggregation is proposed, as shown in the appendix Figure 2 shown, which is mainly divided into multi-scale atrous convolution and channel-spatial aggregation.
[0009] In the multi-scale dilated convolution part, feature extraction for different scales is achieved through five parallel convolutional branches, each branch configured with a different dilation rate: the first branch uses a 1×1 convolutional kernel without changing the spatial scale; the second branch uses a 3×3 convolutional kernel with a dilation rate of 6 to moderately expand the receptive field; the third branch uses a 3×3 convolutional kernel with a dilation rate of 12 to further expand the receptive field to capture more extensive context information; the fourth branch uses a 3×3 convolutional kernel with a dilation rate of 18 to provide the widest receptive field; the fifth branch uses global average pooling to extract global context features and integrate them into local features to enhance the relationship between the local and the whole. The multi-scale dilated convolution increases the receptive field without losing resolution and captures multi-scale context information.
[0010] The five branches are concatenated in the channel dimension to form a comprehensive feature map, and channel calibration and spatial calibration are performed on it. Channel calibration is responsible for evaluating the importance of different channels and strengthening important feature channels. First, global average pooling is performed on the comprehensive feature map to obtain the global features of each channel. Then, the importance weights of different channels are learned through the ReLU activation function and the Sigmoid activation function. Finally, the importance weights are multiplied element-wise with the original feature map to achieve channel weighting. Spatial calibration is responsible for focusing on important regions of the image. Global pooling is performed on the comprehensive feature map in the channel dimension to obtain a spatial feature map. Through a 1×1 convolution and the Sigmoid activation function, the importance weights of different spatial positions are learned. Finally, the importance weights are multiplied element-wise with the original features to achieve spatial weighting. The outputs of channel calibration and spatial calibration are added element-wise to obtain enhanced features, which are further enhanced by adding element-wise to the aggregated features of the five branches. The enhanced features are dimension-reduced through a 1×1 convolution to obtain the final features. Channel-spatial aggregation realizes the dynamic adjustment of the importance weights of features in different dimensions, captures detailed information and global information, and enhances the relationship between parts and the whole.
[0011] Step 2, distributed query mechanism. The planning performance of the current end-to-end work completely depends on the prior information of detection, mapping, and motion prediction. These prior information are the only sources of planning. However, low-quality prior information will inevitably lead to a high planning collision rate and a large L2 error. To decouple their dependence relationship, a distributed query mechanism combined with a distributed decoder is designed. A task-specific learnable feature query is customized for each task. The distributed decoder is used to make the feature query of each module interact with the most original features, ensuring that relevant information of each task is captured from the original input, thereby decoupling the dependence relationship of each task.
[0012] The distributed decoder is as attached Figure 3As shown, it consists of perspective aggregation and a feed-forward neural network, which are distributed in the initialization steps of the scene understanding decoder and the joint probability motion planner. The perspective aggregation therein can efficiently perform feature decoding.
[0013] The distributed queries are distributed in the initialization parts of various tasks in the end-to-end pipeline, as shown in the appendix. Figure 4 Considering the training resource limitations, the occupancy task is not used, but instead, prediction and planning are jointly modeled. The final end-to-end model includes object detection, online mapping, and joint probability motion planning tasks. Among them, object detection and online mapping are auxiliary tasks that supervise the joint probability motion planning, which is the main task, to improve interpretability. In the distributed decoders of the map branch and the detection branch of scene understanding, the learnable initialization map query M0 and the learnable initialization agent query A0 are updated separately from the current scene feature query T: t M
[0014] M t = MHPA(M0, T t , T t )
[0015] A t = MHPA(A0, T t , T t )
[0016] where MHPA(M0, T t , T t ) represents perspective aggregation, which uses M0, T t and T t as the query, key, and value respectively, and M t is the updated map query; MHPA(A0, T t , T t ) represents perspective aggregation, which uses A0, T t and T t as the query, key, and value respectively, and A t is the updated agent query.
[0017] In the distributed decoder of joint probability motion planning, the learnable agent motion query t for planning and the learnable ego-vehicle motion query are obtained from the current scene feature query T
[0018]
[0019] where represents the perspective aggregation mechanism, which uses T t and T tAs query, key, and value, they are distributed agent motion query and distributed ego-vehicle motion query, Q e ″ i which contains the ego-vehicle state.
[0020] The distributed query divides the planning information sources into two paths. The first path is the interaction between the learnable initialization query of a specific task in the distributed decoder and the original features, and the second path is the scene understanding and the prior information stored in the end-to-end memory pool. The two paths of information complement each other to ensure that the planning output will not be negatively affected by the prior information. This makes the inference process more flexible, that is, when interpretability is needed, multi-task modes such as detection and mapping are called, and when real-time performance is needed, the multi-task mode is deactivated and run at a high frame rate, significantly improving the running efficiency and enabling more frequent replanning to dynamically adjust the plan, thereby enhancing the safety of deployment.
[0021] Step 3: Scene understanding. Scene understanding is divided into online map and object detection. As shown in the appendix Figure 5 it provides prior information for the joint probabilistic motion planning and supervises the learning of the end-to-end network as an auxiliary task to enhance interpretability. The results of scene understanding are respectively characterized by the object instance bounding boxes and corresponding attributes, and the map elements and corresponding attributes.
[0022] (1) Online map.
[0023] Most current end-to-end works represent static scenes based on a deterministic paradigm, which will result in the static map elements followed by the planning being deterministic, making any map errors (e.g., offset or misplacement of map elements) cause deviations in downstream behaviors. To address this, first, the map decoder decodes the map query into map element points in space, then probabilistically models the positions and their categories of the static map elements based on the Laplace distribution, and passes them to the planning in the form of queries, enabling the planning to fully consider the errors in the static representation and reduce this uncertainty. Consider three types of map elements: lane dividers, road boundaries, and crosswalks. The map branch consists of a distributed decoder and a map decoder. The map decoder consists of cross-attention, self-attention, perspective aggregation mechanism, feed-forward network, and an uncertainty regression head and classification head based on the Laplace probability distribution.
[0024] The distributed decoder outputs M t and combines the historical map queries in the streaming memory pool with M tAggregation is performed and then it enters the map decoder. The streaming memory pool stores historical feature queries with high confidence. The streaming query is used as a medium for historical feature transmission to transmit information frame by frame, thus avoiding interaction with all frame images and reducing the computational cost while maintaining high performance. The streaming memory pool is end-to-end and contains all historical queries for scene understanding and joint probabilistic motion planning.
[0025] In the cross-attention mechanism, the map query interacts with the current frame and historical frames:
[0026] M′ = MHCA(M t +M t-k ,M t-k ,M t-k )
[0027] where MHCA(M t +M t-k ,M t-k ,M t-k ) represents the cross-attention mechanism, using M t +M t-k , M t-k and M t-k as the query, key, and value respectively. M′ is the updated query, M t is the updated map query in the distributed decoder, and M t-k is the historical map query, which consists of map queries with high confidence in the previous frames.
[0028] Then, interactions within the current frame are performed in the self-attention mechanism:
[0029] M″ = MHSA(M′, M′, M′)
[0030] where MHSA(M′, M′, M′) represents the self-attention mechanism, using M′, M′, and M′ as the query, key, and value respectively, and M″ is the updated query.
[0031] In the perspective aggregation mechanism, the current scene feature T t is decoded:
[0032] M = MHPA(M″, T t , T t )
[0033] where MHPA(M″, T t , T t ) represents the perspective aggregation mechanism, using M″, T t and T t as the query, key, and value respectively, and M is the updated map query.
[0034] After the interaction, the positions and categories of static map elements are regressed and recognized respectively. There are many sources of uncertainty in static scene modeling methods, such as the transformation from PV to BEV, the positions of polyline vertices, the connection of map elements, and occlusion, etc. Almost all current studies use point regression and classification heads to predict the positions of polyline vertices and identify the types of elements in the static scenes they belong to respectively. Therefore, in order to enhance the generalization of the proposed static uncertainty method, based on the two general output structures of the original point regression and classification heads, regression probability and category probability are added, so that traffic participants can plan their driving by following more reasonable static probability map elements.
[0035] First is the regression probability. The regression head usually adopts a simple MLP architecture. For each map element, the regression head generates a two-dimensional vector to represent the point coordinates (x, y) in the bird's-eye view. To convert it into a probability model with uncertainty, the present invention uses a regression head that can output uncertainty parameters related to the predicted points. The Laplace distribution has a sharp peak and heavier tails, and is particularly good at handling outlier anomalies to model various sources of uncertainty. Therefore, the present invention uses the Laplace distribution to model each vertex p=(p1, p2) of the map element. Correspondingly, for a map element E with P vertices (denoted as where a represents the a-th vertex, and its joint probability density distribution is
[0036]
[0037] where and are the location parameter and scale parameter of the Laplace distribution of the b-th dimension of the a-th vertex of the map element E.
[0038] Second is the classification probability. The classification head outputs a class confidence score for each regression vertex, and these scores (logits) provide the classification distribution in the form of a probability distribution. This distribution already contains sufficient semantic information, so the probability distribution of these classification scores (logits) can be directly passed to the planning model without additional processing steps.
[0039] To encode the static uncertainty into the joint probability motion planning to reduce this uncertainty error in planning, for the Laplace distribution of the a-th vertex of the map element E, the map point location parameter λ, the uncertainty scale parameter s, and the category parameter l are concatenated, and then encoded using a multi-layer perceptron MLP:
[0040]
[0041] The obtained unc p fuses the probability map vertex features.
[0042] At this time, the updated map query M has static uncertainty, which is used as prior information for planning on the one hand and to update the streaming memory pool for future use on the other hand.
[0043] (2) Object detection.
[0044] The object detection branch is similar to the online map branch and needs to combine and interact the A obtained by the distributed decoder t with the historical detection queries in the streaming memory pool. Since the pose coordinates of all traffic participants will change before and after movement, the MLN (Motion aware Layer Normalization) is used to perform pose coordinate transformation on the historical detection queries and A t This invention first updates the learnable initial agent query A0 based on the distributed decoder and the detection decoder from the scene feature query T t so as to obtain the updated agent query A:
[0045] A′ = MHCA(MLN(A t + A t-k ), MLN(A t-k ), MLN(A t-k ))
[0046] A″ = MHSA(A′, A′, A′)
[0047] A = MHPA(A″, T t , T t )
[0048] Finally, a 3D object detection head is used to decode the relevant information of each agent.
[0049] Step four, joint probability motion planning. Existing methods assume a deterministic relationship between the actions of the environment and all traffic participants during planning and ignore the influence of the environment on the ego-vehicle planning and the influence of the ego-vehicle on other traffic participants, which will increase the uncertainty, resulting in a high collision rate and a large L2 error. To solve this problem, this invention fully considers the high-order interaction among the ego-vehicle, agents, static map, and streaming temporal information, and simultaneously models the agent motion prediction and the ego-vehicle trajectory planning as joint probability motion planning, representing this probability distribution with planning vocabulary and probability fields. The joint probability motion planning consists of a distributed decoder and a planning decoder, as shown in the appendix Figure 6 as shown.
[0050] The planning actions are represented as a probability distribution of driving samples, and corresponding actions are sampled from the distribution at each time step to control the vehicle. The planning actions are high-dimensional and continuous in time and space. The planning actions are discretized into planning vocabulary clusters And sample N representative planning vocabulary from it. Each action in the planning vocabulary is represented as a trajectory sequence e = (x1, y1, x2, y2,..., x T , y T ), and each trajectory corresponds to a future timestamp. Since the actions are continuous on the time axis, the probability corresponding to the action at each moment is continuous. Then, use probability field modeling to model the continuous mapping from the planning action to the probability distribution . Encode the environmental information as Q env , and encode the planning action as the initial ego-vehicle motion query . After interaction, obtain the final ego-vehicle motion planning query ego-vehicle query and the environmental information Q env Perform high-order interaction based on Transformer:
[0051]
[0052] Among them, the trajectory e consists of coordinate values (trajectory sequence) e = (x1, y1, x2, y2,..., x T , y T ), Q hyb is a mixed instance motion query composed of the agent motion query and the ego-vehicle motion query, is the historical mixed instance motion query. The environmental information Q env includes the mixed instance motion query Q hyb composed of the ego-vehicle motion query and the agent motion query, the memory pool mixed instance motion query detection query A and map query M.
[0053] The input Q hyb of the joint probability motion planning has two sources. Given the initial ego-vehicle motion query Use the historical detection query, the current detection query, and the current map query as the initial agent motion query The above two queries are the first source. This can use the multi-modal mixed instance motion query composed of the agent and the ego-vehicle as a medium to aggregate rich semantic information from the static map and the dynamic agent, and benefit from the prior of scene understanding, including information such as static uncertainty. Obtain the learnable agent motion query and the learnable ego-vehicle motion query These two are the second source of planning. This can obtain the ego-vehicle and agent motion information from the original scene features, ignoring the interference of redundant noise. Finally, combine the queries from the two sources to obtain the input of the motion planning.
[0054] Initial self-vehicle motion query from the first source Expressed as:
[0055]
[0056] where (x1, y1)...(x T , y T ) represent the coordinate position loc, T is the coordinate index, En() is the position encoding function, En() maps each coordinate loc to a high-dimensional embedding space and is applied to each coordinate value of the trajectory e respectively,
[0057] En(loc) = concat[ω(loc, 0), ω(loc, 1),..., ω(loc, J - 1)]
[0058] Inspired by the sine position encoding of the Transformer mechanism, the position encoding function is defined as follows:
[0059] ω(loc, j) = concat[cos(loc / 1 - 00 - 2πj / J , sin(loc / 1 - 00 - 2πj / J ))]
[0060] where j represents the encoding dimension index, J is the encoding dimension, ω is the sine-cosine position encoding function that converts the trajectory point coordinates into high-dimensional features, and concat[·] represents concatenation. These functions are used to map the continuous input coordinates to a higher-dimensional space to better approximate higher-frequency field functions, enabling the model to perceive the relative positions of the trajectory points and avoid learning biases in absolute positions.
[0061] Combine the current scene understanding information with the prior information in the streaming memory pool to obtain the initial agent motion query from the first source
[0062]
[0063] where MHCA represents the cross-attention mechanism and MLP represents the multi-layer perceptron, represents the agent's historical position information, A m represents the historical detection query, A represents the current detection query, M represents the current map query, represents the agent's learned query embedding.
[0064] Aggregate the first source and the second source to obtain the agent motion query and the self-vehicle motion query
[0065]
[0066] First, the agent motion query and the ego-vehicle motion query are aggregated to obtain the current hybrid instance motion query Q hyb :
[0067]
[0068] To fully consider the mutual influence between the agent and the ego-vehicle, and between agents, the interaction within the current hybrid instance motion query is carried out based on the self-attention mechanism:
[0069] Q hyb = MHSA(Q hyb , Q hyb , Q hyb )
[0070] where MHSA(Q hyb , Q hyb , Q hyb ) is multi-head self-attention, using Q hyb , Q hyb and Q hyb as the query, key, and value respectively.
[0071] To fully consider the ego-vehicle historical information and agent historical information, the current hybrid instance motion query is interacted with the historical hybrid instance motion query:
[0072]
[0073] where is cross-attention, using Q hyb , and as the query, key, and value respectively.
[0074] To capture the current agent position and attributes, the current hybrid instance motion query is interacted with the current detection query:
[0075] Q hyb = MHCA(Q hyb , A, A)
[0076] To achieve accurate prediction and planning, both the agent and the ego-vehicle need to consider the high-level semantic information containing the static uncertainty map, and the current hybrid instance motion query is interacted with the online map query:
[0077] Q hyb = MHCA(Q hyb , M, M)
[0078] where MHCA(Q hyb,M,M) are multi-head cross attention, using Q hyb , M and M as query, key and value. Q hyb is the updated current mixed instance motion query, including the current agent motion query Q a And self-driving sports query Q e .
[0079] The learned mixed instance motion query integrates high-order vehicle-environment-time interaction relationships, including static representation of uncertainty and dynamic representation of uncertainty. Finally, the corresponding probabilistic planning results are calculated by MLP:
[0080]
[0081] The maximum probability trajectory e*=argmaxp(e) is used as the current trajectory among the multimodal trajectories at different times to reduce static uncertainty and dynamic uncertainty. At the same time, the updated mixed instance motion query is used to update the streaming memory pool to enter the planning of the next frame.
[0082] Step six: end-to-end learning.
[0083] The end-to-end model in this invention includes three task loss functions, namely, detection loss L det , map loss L map and the joint probabilistic motion planning loss L motion_plan . Joint probabilistic motion planning is used as the main task, detection and mapping are used as auxiliary tasks, and the main task supervises and improves end-to-end interpretability. The loss function is calculated as follows:
[0084] L=λ1L det +λ2L map +λ3L motion_plan
[0085] Among them, λ1, λ2, and λ3 are the corresponding loss weights.
[0086] In the target detection task, the present invention uses the Hungarian algorithm to match the true value and the predicted value, and the detection loss L det is the classification loss L FocalLoss and the regression loss L L1 A linear combination of:
[0087] L det_cls =L FocalLoss ,L det_reg =L L1
[0088] L det =L det_cls +L det_reg
[0089] The map loss consists of the map element classification loss L Focal and the negative log-likelihood loss L NLL for map instance regression:
[0090] L map_cls = L Focal , L map_reg = L NLL
[0091] L map = L map_cls + L map_reg
[0092] In joint probabilistic motion planning, the ground truth trajectory is added to the planning vocabulary as a positive sample, and the remaining trajectories are negative samples. The KL divergence is used to calculate the distribution loss between the predicted distribution and the ground truth distribution The present invention also adds a self-vehicle-agent collision constraint L ea , a self-vehicle-road boundary crossing constraint L eb , and a self-vehicle-lane direction constraint L el :
[0093]
[0094] Based on the above method for constructing an uncertainty end-to-end autonomous driving model that combines distributed queries and fuses spatio-temporal information, the present invention also proposes a corresponding model, which is the model built and trained in the above steps one to five.
[0095] The present invention also proposes a vehicle device in which the above model is arranged.
[0096] The beneficial effects of the present invention are as follows:
[0097] (1) The present invention proposes an uncertainty end-to-end autonomous driving that combines distributed queries and fuses spatio-temporal information. At the same time, by using the end-to-end streaming temporal strategy and the sparse representation method, it solves the safety problems caused by high collision rates and large L2 errors in the deterministic paradigm and improves resource efficiency.
[0098] (2) The present invention proposes to model static uncertainty and dynamic uncertainty simultaneously in the form of probability distributions, jointly model agent motion prediction and trajectory planning, and perform high-order interaction between the self-vehicle and the environment to reduce the collision rate and L2 error caused by static uncertainty and dynamic uncertainty.
[0099] (3) The present invention proposes multi-scale hole aggregation, which highlights the contribution of important features and expands the receptive field, and solves the problem of dimensional contribution in the deterministic case.
[0100] (4) The present invention proposes a distributed query, which ensures that the plan has multiple reliable information sources and solves the problem of the plan's dependence on upstream tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0101] Figure 1 is a schematic diagram of the overall structure of an end-to-end autonomous driving system;
[0102] Figure 2 is a schematic diagram of multi-scale hole aggregation;
[0103] Figure 3 is a schematic diagram of a distributed decoder;
[0104] Figure 4 is a schematic diagram of a distributed query mechanism;
[0105] Figure 5 is a schematic diagram of scene understanding;
[0106] Figure 6 is a schematic diagram of joint probability motion planning; DETAILED DESCRIPTION OF THE INVENTION
[0107] The present invention will be further described below in conjunction with the drawings and the detailed implementation, but the protection scope of the present invention is not limited thereto.
[0108] The detailed implementation of the present invention is as follows:
[0109] (1) Multi-scale hole aggregation. First, use a convolutional neural network to extract image features, and then use multi-scale hole aggregation to enhance the features to obtain the final scene features. Different dimensions of the convolutional neural network represent different features or patterns, and the importance of the information contained is different. However, when extracting features, it is default that the contribution importance of different dimensions is definitely the same, resulting in unreasonable weight allocation. In addition, the dynamic traffic participants in the driving scene are very different in size, shape and style from the static map, and the fixed receptive field of the convolutional neural network limits its ability to capture features at different scales. To address this problem, multi-scale hole aggregation is proposed, as shown in the attached Figure 2 figures, which is mainly divided into multi-scale hole convolution and channel-spatial aggregation.
[0110] In the multi-scale dilated convolution part, feature extraction for different scales is achieved through five parallel convolutional branches, each branch configured with a different dilation rate: the first branch uses a 1×1 convolutional kernel without changing the spatial scale; the second branch uses a 3×3 convolutional kernel with a dilation rate of 6 to moderately expand the receptive field; the third branch uses a 3×3 convolutional kernel with a dilation rate of 12 to further expand the receptive field to capture more extensive context information; the fourth branch uses a 3×3 convolutional kernel with a dilation rate of 18 to provide the widest receptive field; the fifth branch uses global average pooling to extract global context features and integrate them into local features to enhance the relationship between the local and the whole. The multi-scale dilated convolution increases the receptive field without losing resolution and captures multi-scale context information.
[0111] The five branches are concatenated in the channel dimension into a comprehensive feature map and channel calibration and spatial calibration are performed on it. Channel calibration is responsible for evaluating the importance of different channels and strengthening important feature channels. First, global average pooling is performed on the comprehensive feature map to obtain the global features of each channel. Then, the importance weights of different channels are learned through the ReLU activation function and the Sigmoid activation function. Finally, the importance weights are multiplied element-wise with the original feature map to achieve channel weighting. Spatial calibration is responsible for focusing on important regions of the image. Global pooling is performed on the comprehensive feature map in the channel dimension to obtain the spatial feature map. Through a 1×1 convolution and the Sigmoid activation function, the importance weights of different spatial positions are learned. Finally, the importance weights are multiplied element-wise with the original features to achieve spatial weighting. The outputs of channel calibration and spatial calibration are added element-wise to obtain the enhanced features, which are further enhanced by adding element-wise to the aggregated features of the five branches. The enhanced features are dimensionally reduced through a 1×1 convolution to obtain the final features. Channel-spatial aggregation realizes dynamic adjustment of the importance weights of features in different dimensions, captures detailed information and global information, and enhances the relationship between parts and the whole.
[0112] (2) Distributed query mechanism. Currently, the planning performance of the end-to-end work completely depends on the prior information of detection, mapping, and motion prediction, which is the only source of planning. However, low-quality prior information will inevitably lead to a high planning collision rate and a large L2 error. To decouple their dependency relationship, a distributed query mechanism combined with a distributed decoder is designed, which customizes a task-specific learnable feature query for each task. The distributed decoder enables the feature query of each module to interact with the most original features, ensuring the capture of relevant information for each task from the original input, thereby decoupling the dependency relationship between each task.
[0113] The distributed decoder is as shown in the appendix Figure 3As shown, it consists of perspective aggregation and a feed-forward neural network, which are distributed in the initialization of the scene understanding decoder and the joint probability motion planner. The perspective aggregation therein can efficiently perform feature decoding.
[0114] The distributed queries are distributed in the initialization parts of various tasks in the end-to-end pipeline, as shown in the appendix. Figure 4 Considering the training resource limitations, the occupancy task is not used, but instead, prediction and planning are jointly modeled. The final end-to-end model includes object detection, online mapping, and joint probability motion planning tasks. Among them, object detection and online mapping are auxiliary tasks that supervise the joint probability motion planning, which is the main task, to improve interpretability. In the distributed decoders of the map branch and the detection branch of scene understanding, the current scene feature query T t is used to separately update the learnable initialization map query M0 and the learnable initialization agent query A0:
[0115] M t = MHPA(M0, T t , T t )
[0116] A t = MHPA(A0, T t , T t )
[0117] where MHPA(M0, T t , T t ) represents perspective aggregation, which uses M0, T t and T t as the query, key, and value respectively, and M t is the updated map query; MHPA(A0, T t , T t ) represents perspective aggregation, which uses A0, T t and T t as the query, key, and value respectively, and A t is the updated agent query.
[0118] In the distributed decoder of joint probability motion planning, the current scene feature query T t is used to obtain the learnable agent motion query for planning and the learnable ego-vehicle motion query
[0119]
[0120] where represents the perspective aggregation mechanism, which uses T t and T tAs queries, keys, and values, they are distributed agent motion queries and distributed ego-vehicle motion queries respectively, which contain the ego-vehicle state.
[0121] The distributed query divides the sources of planning information into two paths. The first path is the interaction between the learnable initialization query of a specific task in the distributed decoder and the original features, and the second path is scene understanding and the prior information stored in the end-to-end memory pool. The two paths of information complement each other to ensure that the planning output is not negatively affected by the prior information. This makes the inference process more flexible, that is, calling multitask modes such as detection and mapping when interpretability is needed, and deactivating the multitask mode when real-time performance is needed to run at a high frame rate, significantly improving the running efficiency, enabling more frequent replanning to dynamically adjust the plan, and thus improving the safety of deployment.
[0122] (3) Scene understanding. Scene understanding is divided into online map and object detection. As shown in the appendix Figure 5 it provides prior information for joint probability motion planning and supervises the learning of the end-to-end network as an auxiliary task to enhance interpretability. The results of scene understanding are represented by the object instance bounding boxes and corresponding attributes, as well as the map elements and corresponding attributes respectively.
[0123] 1) Online map.
[0124] Most current end-to-end work represents static scenes based on a deterministic paradigm, which will result in the static map elements followed by the planning being deterministic, making any map errors (e.g., offsets or misplacements of map elements) cause downstream behavior deviations. To address this, first use the map decoder to decode the map query into map element points in space, then probabilistically model the positions and categories of static map elements based on the Laplace distribution, and transfer them to the planning in the form of queries, enabling the planning to fully consider the errors of the static representation and reduce this uncertainty. Consider three types of map elements: lane dividers, road boundaries, and crosswalks. The map branch consists of a distributed decoder and a map decoder. The map decoder consists of cross-attention, self-attention, perspective aggregation mechanism, feed-forward network, and an uncertainty regression head and classification head based on the Laplace probability distribution.
[0125] The distributed decoder outputs M t , aggregates the historical map queries in the streaming memory pool with M t , and then enters the map decoder. The streaming memory pool stores historical feature queries with high confidence. Using the streaming query as a medium for historical feature transmission to transmit information frame by frame, thus avoiding interacting with all frame images and reducing the computational cost while maintaining high performance. The streaming memory pool is end-to-end and contains all historical queries of scene understanding and joint probability motion planning.
[0126] In the cross-attention mechanism, the map query interacts with the current frame and historical frames:
[0127] M′ = MHCA(M t +M t-k ,M t-k ,M t-k )
[0128] where MHCA(M t +M t-k ,M t-k ,M t-k ) represents the cross-attention mechanism, using M t +M t-k , M t-k and M t-k as the query, key, and value respectively. M′ is the updated query, M t is the updated map query in the distributed decoder, and M t-k is the historical map query, which consists of map queries with high confidence in previous frames.
[0129] Then, interactions within the current frame are performed in the self-attention mechanism:
[0130] M″ = MHSA(M′, M′, M′)
[0131] where MHSA(M′, M′, M′) represents the self-attention mechanism, using M′, M′, and M′ as the query, key, and value respectively, and M″ is the updated query.
[0132] In the perspective aggregation mechanism, the current scene feature T t is decoded:
[0133] M = MHPA(M″, T t , T t )
[0134] where MHPA(M″, T t , T t ) represents the perspective aggregation mechanism, using M″, T t and T t as the query, key, and value respectively, and M is the updated map query.
[0135] After the interaction, the positions and categories of the static map elements are regressed and recognized respectively. There are many sources of uncertainty in the static scene modeling method, such as the transformation from PV to BEV, the positions of the polyline vertices, the connection of map elements, and occlusion, etc. Almost all current studies use point regression and classification heads to predict the positions of the polyline vertices and identify the types of elements belonging to the static scene respectively. Therefore, in order to enhance the generalization of the proposed static uncertainty method, on the basis of the two general output structures of the original point regression and classification heads, regression probability and category probability are added, so that traffic participants can plan their driving by following a more reasonable static probability map element.
[0136] First is the regression probability. The regression head usually adopts a simple MLP architecture. For each map element, the regression head generates a two-dimensional vector to represent the point coordinates (x, y) in the bird's-eye view. In order to convert it into a probability model with uncertainty, the present invention uses a regression head that can output uncertainty parameters related to the predicted points. The Laplace distribution has a sharp peak and heavier tails, and is particularly good at handling outlier outliers to model various sources of uncertainty. Therefore, the present invention uses the Laplace distribution to model each vertex p=(p1,p2) of the map element. Correspondingly, the map element E with P vertices (denoted as where a represents the a-th vertex, and its joint probability density distribution is
[0137]
[0138] where and are the location parameter and scale parameter of the Laplace distribution of the b-th dimension of the a-th vertex of the map element E.
[0139] Second is the classification probability. The classification head outputs a class confidence score for each regression vertex, and these scores (logits) provide the classification distribution in the form of a probability distribution. This distribution already contains sufficient semantic information, so the probability distribution of these classification scores (logits) can be directly passed to the planning model without additional processing steps.
[0140] In order to encode the static uncertainty into the joint probability motion planning to reduce this uncertainty error in the planning, for the Laplace distribution of the a-th vertex of the map element E, the map point position parameter λ, the uncertainty scale parameter s, and the category parameter l are concatenated, and then a multi-layer perceptron MLP is used to encode it:
[0141]
[0142] The obtained unc p fuses the probability map vertex features.
[0143] At this time, the updated map query M has static uncertainty, which is used as prior information for planning on the one hand and to update the streaming memory pool for future use on the other hand.
[0144] 2) Object detection.
[0145] The object detection branch is similar to the online map branch and needs to combine and interact the A obtained by the distributed decoder t with the historical detection queries in the streaming memory pool. Since the pose coordinates of all traffic participants change before and after movement, the MLN is used to perform pose coordinate transformation on the historical detection queries and A t We first update the learnable initial agent query A0 based on the distributed decoder and the detection decoder from the scene feature query T t to obtain the updated agent query A:
[0146] A′ = MHCA(MLN(A t + A t-k ), MLN(A t-k ), MLN(A t-k ))
[0147] A″ = MHSA(A′, A′, A′)
[0148] A = MHPA(A″, T t , T t )
[0149] Finally, a 3D object detection head is used to decode the relevant information of each agent.
[0150] (4) Joint probabilistic motion planning. Existing methods assume a deterministic relationship between the actions of the environment and all traffic participants during planning and ignore the impact of the environment on the ego-vehicle's planning and the impact of the ego-vehicle on other traffic participants, which will increase uncertainty, resulting in a high collision rate and a large L2 error. To solve this problem, we fully consider the high-order interaction between the ego-vehicle, agents, static map, and streaming temporal information, and simultaneously model the agent motion prediction and the ego-vehicle trajectory planning as joint probabilistic motion planning, representing this probability distribution with planning vocabulary and probability fields. The joint probabilistic motion planning consists of a distributed decoder and a planning decoder, as shown in the appendix Figure 6 as shown.
[0151] Represent the planned actions as a probability distribution of driving samples, and sample the corresponding actions from the distribution at each time step to control the vehicle. The planned actions are high-dimensional and continuous in time and space. Discretize the planned actions into planning vocabulary clusters And sample N representative planning words from it. Each action in the planning words is represented as a trajectory sequence e = (x1, y1, x2, y2,..., x T , y T ), and each trajectory corresponds to a future timestamp. Since the actions are continuous on the time axis, the probabilities corresponding to the actions at each moment are continuous. Then, use probability field modeling to model the continuous mapping from the planning actions to the probability distribution . Encode the environmental information as Q env , and encode the planning actions as the initial ego-vehicle motion query . After interaction, obtain the final ego-vehicle motion planning query The ego-vehicle query and the environmental information Q env Perform high-order interaction based on Transformer:
[0152]
[0153] Among them, the trajectory e is the coordinate value (trajectory sequence) e = (x1, y1, x2, y2,..., x T , y T ), Q hyb is a mixed instance motion query composed of the agent motion query and the ego-vehicle motion query, is the historical mixed instance motion query. The environmental information Q env includes the mixed instance motion query Q hyb composed of the ego-vehicle motion query and the agent motion query, the memory pool mixed instance motion query the detection query A and the map query M.
[0154] The input Q hyb of the joint probability motion planning has two sources. Given the initial ego-vehicle motion query Use the historical detection query, the current detection query, and the current map query as the initial agent motion query The above two queries are the first source. This can use the multi-modal mixed instance motion query composed of the agent and the ego-vehicle as a medium to aggregate rich semantic information from the static map and the dynamic agent, benefit from the prior of scene understanding, and contain information such as static uncertainty. The learnable agent motion query and the learnable ego-vehicle motion query are obtained through the distributed decoder. These two are the second source of the planning. This can obtain the ego-vehicle and agent motion information from the original scene features and ignore the interference of redundant noise. Finally, combine the queries from the two sources to obtain the input of the motion planning.
[0155] The initial ego-vehicle motion query of the first source Expressed as:
[0156]
[0157] where (x T , y T ) represents the coordinate position loc, T is the coordinate index, En() is the position encoding function, En() maps each coordinate loc to a high-dimensional embedding space and is applied to each coordinate value of the trajectory e respectively,
[0158] En(loc) = concat[ω(loc, -), ω(loc, 1),..., ω(loc, J - 1)]
[0159] Inspired by the sine position encoding of the Transformer mechanism, the position encoding function is defined as follows:
[0160] ω(loc, j) = concat[cos(loc / 10000 2πj / J , sin(loc / 10000 2πj / J ))]
[0161] where j represents the encoding dimension index, J is the encoding dimension, ω is the sine-cosine position encoding function that converts the trajectory point coordinates into high-dimensional features, and concat[·] represents concatenation. Using these functions, the continuous input coordinates are mapped to a higher-dimensional space to better approximate the higher-frequency field function, enabling the model to perceive the relative positions of the trajectory points and avoid learning biases of absolute positions.
[0162] Combine the current scene understanding information with the prior information in the streaming memory pool to obtain the initialization of the first source of the agent's motion query
[0163]
[0164] where MHCA represents the cross-attention mechanism and MLP represents the multi-layer perceptron, represents the agent's historical position information, A m represents the historical detection query, A represents the current detection query, M represents the current map query, represents the agent's learning query embedding.
[0165] Aggregate the first source and the second source to obtain the agent's motion query and the ego-vehicle motion query
[0166]
[0167] Here, first the agent's motion query and ego-vehicle motion query Aggregate to obtain the current hybrid instance motion query Q hyb :
[0168]
[0169] To fully consider the mutual influence between the agent and the ego-vehicle, and between agents, perform the interaction within the current hybrid instance motion query based on the self-attention mechanism:
[0170] Q hyb = MHSA(Q hyb , Q hyb , Q hyb )
[0171] where MHSA(Q hyb , Q hyb , Q hyb ) is multi-head self-attention, using Q hyb , Q hyb and Q hyb as query, key, and value respectively.
[0172] To fully consider the ego-vehicle historical information and agent historical information, perform the interaction between the current hybrid instance motion query and the historical hybrid instance motion query:
[0173]
[0174] where, is cross-attention, using Q hyb , and as query, key, and value respectively.
[0175] To capture the current agent position and attributes, perform the interaction between the current hybrid instance motion query and the current detection query:
[0176] Q hyb = MHCA(Q hyb , A, A)
[0177] To achieve accurate prediction and planning, both the agent and the ego-vehicle need to consider the high-level semantic information containing the static uncertainty map, and perform the interaction between the current hybrid instance motion query and the online map query:
[0178] Q hyb = MHCA(Q hyb , M, M)
[0179] where MHCA(Q hyb , M, M) is multi-head cross-attention, using Q hyb, M and M are used as query, key, and value. Q hyb is the updated current hybrid instance motion query, including the current agent motion query Q a and the ego - vehicle motion query Q e .
[0180] The learned hybrid instance motion query fuses high - order ego - vehicle - environment - temporal interaction relationships, including uncertain static representations and uncertain dynamic representations. Finally, the corresponding probabilistic planning results are calculated through MLP:
[0181]
[0182] Among the multi - modal trajectories at different times, the maximum - probability trajectory e* = argmax p(e) is adopted as the current trajectory to reduce static uncertainty and dynamic uncertainty. Meanwhile, the updated hybrid instance motion query is used to update the streaming memory pool to enter the planning of the next frame.
[0183] (5) End - to - end learning.
[0184] The end - to - end model in the present invention includes three task loss functions, namely detection, map, and joint probabilistic motion planning. Joint probabilistic motion planning is the main task, and detection and map are auxiliary tasks. The main task supervises and improves the interpretability of the end - to - end. The loss function is calculated as follows:
[0185] L = λ1L det +λ2L map +λ3L motion_plan
[0186] where λ1, λ2, and λ3 are loss weights.
[0187] In the object - detection task, the present invention uses the Hungarian algorithm to match the ground truth and predicted values. The detection loss L det is the linear combination of the classification loss L FocalLoss and the regression loss L L1 :
[0188] L det_cls = L FocalLoss , L det_reg = L L1
[0189] L det = L det_cls +L det_reg
[0190] The map loss is composed of the map - element classification loss L Focal and the negative log - likelihood loss L NLL of the map - instance regression:
[0191] Lmap_cls = L Focal , L map_reg = L NLL
[0192] L map = L map_cls + L map_reg
[0193] In joint probability motion planning, the true trajectory is added to the planning vocabulary as a positive sample, and the remaining trajectories are negative samples. The KL divergence is used to calculate the distribution loss between the predicted distribution and the true distribution. The present invention also adds a self-vehicle-agent collision constraint L ea , a self-vehicle-road boundary crossing constraint L eb , a self-vehicle-lane direction constraint L el :
[0194]
[0195] (6) An embodiment of the present invention also proposes an end-to-end autonomous driving device, which uses the built model in combination with devices such as on-vehicle cameras to implement the end-to-end autonomous driving method that combines distributed queries and fuses the uncertainty of spatio-temporal information.
[0196] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation manners of the present invention, and they are not intended to limit the protection scope of the present invention. Any equivalent manners or changes that do not depart from the technology created by the present invention should be included in the protection scope of the present invention.
Claims
1. A method for constructing an end-to-end autonomous driving model that combines distributed query and fuses the uncertainty of spatio-temporal information, characterized in that The following are included: S1. Design multi-scale dilated convolution aggregation, including: S1.
1. Use a convolutional neural network to extract image features; S1.
2. Use multi-scale dilated aggregation to enhance the features to obtain the final scene features; multi-scale dilated convolution increases the receptive field without losing the resolution and captures multi-scale context information; S2. Design distributed queries, customize a set of specific learnable query features for each task of autonomous driving, decouple the dependency relationship between the planning task and the upstream task, enable the query features of each task to interact with the obtained original features, ensure the capture of relevant information for each task from the original input, and improve the planning stability; S3. Design a scene understanding module, including: S3.
1. Perform spatio-temporal interaction between different queries and corresponding features and decode the high-dimensional features to obtain object detection and online map results; S3.
2. Use the obtained high-confidence perception queries to update the streaming memory pool, enter the next frame of perception, and provide prior information for joint probabilistic motion planning, and supervise the learning of the end-to-end network as an auxiliary task to enhance interpretability; the results of scene understanding are respectively characterized by object instance bounding boxes and corresponding attributes, and map elements and corresponding attributes; S4. Design a joint probabilistic motion planning module, simultaneously model the agent motion prediction and the ego vehicle trajectory planning as joint probabilistic motion planning, and fully consider the high-order interaction between the ego vehicle, the agent, the map elements, and the streaming temporal information.
2. The method for constructing an autonomous driving model according to claim 1, wherein The implementation of S1.2 includes: S1.2.
1. Implement feature extraction of different scales through five parallel convolutional branches, each branch configured with a different dilation rate: the first branch uses a 1×1 convolutional kernel without changing the spatial scale; the second branch uses a 3×3 convolutional kernel with a dilation rate of 6 to moderately expand the receptive field; the third branch uses a 3×3 convolutional kernel with a dilation rate of 12 to further expand the receptive field to capture a wider context information; the fourth branch uses a 3×3 convolutional kernel with a dilation rate of 18 to provide the widest receptive field; the fifth branch uses global average pooling to extract global context features and integrate them into local features to enhance the relationship between the local and the whole; S1.2.
2. Concatenate the five branches into a comprehensive feature map in the channel dimension and perform channel calibration and spatial calibration; Channel calibration is responsible for evaluating the importance of different channels and strengthening the important feature channels; first, globally average pool the comprehensive feature map to obtain the global features of each channel; then, learn the importance weights of different channels through the ReLU activation function and the Sigmoid activation function; finally, multiply the importance weights with the original feature map channel by channel to achieve channel weighting; Spatial calibration is responsible for focusing on the important regions of the image, globally pool the comprehensive feature map in the channel dimension to obtain the spatial feature map, learn the importance weights of different spatial positions through a 1×1 convolution and the Sigmoid activation function; finally, multiply the importance weights with the original features element by element to achieve spatial weighting; The output of S1.2.3 channel calibration and spatial calibration is enhanced by element-wise summation to obtain an enhanced feature, which is further enhanced by element-wise summation with the aggregated features of five branches. The enhanced feature is dimension-reduced by a 1×1 convolution to obtain the final feature.
3. The method for constructing an autonomous driving model according to claim 1, wherein, The distributed query of S2 adopts a distributed query mechanism combined with a distributed decoder, customizes learnable feature queries specific to each task, and uses the distributed decoder to interact the feature queries of each part with the most primitive features, ensuring that relevant information for each task is captured from the original input, thereby decoupling the dependencies of each task. The distributed decoder consists of perspective aggregation and a feed-forward neural network, which are distributed in the decoder of the scene understanding module and the initialization link of the joint probability motion planning module. The perspective aggregation therein can efficiently perform feature decoding. The distributed query is distributed in each task initialization part of the end-to-end pipeline: In the distributed decoder of the online map branch and the target detection branch of the scene understanding module, query T from the current scene features t Update the learnable initial map query M0 and the learnable initial agent query A0 respectively: M t = MHPA(M0, T t , T t ) A t = MHPA(A0, T t , T t ) Among them, MHPA(M0, T t , T t ) represents perspective aggregation, using M0, T t and T t as the query, key, and value respectively, and M t is the updated map query; MHPA(A0, T t , T t ) represents perspective aggregation, using A0, T t and T t as the query, key, and value respectively, and A t is the updated agent query; In the distributed decoder for joint probabilistic motion planning, query the current scene features T t Obtain the learnable agent motion query for planning and the learnable ego-vehicle motion query Among them, MHPA(Q, K, V) represents the perspective aggregation mechanism, using Q, K, and V as the query, key, and value respectively, which are the distributed agent motion query and the distributed ego-vehicle motion query respectively, and includes the ego-vehicle state; The distributed query divides the source of planning information into two paths. The first path is the interaction between the learnable initialization query specific to the task in the distributed decoder and the original feature, and the second path is the prior information in the scene understanding and stored in the end-to-end memory pool. The two paths of information complement each other to ensure that the planning output is not negatively affected by the prior information.
4. The method for constructing an autonomous driving model according to claim 1, wherein The implementation of the S3.1 online map branch includes: First, use the map decoder to decode the map query into map element points in space, then probabilistically model the positions and categories of static map elements based on the Laplace distribution and pass them to the planning in the form of a query, enabling the planning to fully consider the error of the static representation and reduce this uncertainty. Considering three types of map elements, namely lane dividers, road boundaries, and crosswalks, the map branch includes a distributed decoder and a map decoder. The map decoder consists of cross-attention, self-attention, a perspective aggregation mechanism, a feed-forward network, and an uncertainty regression head and a classification head based on the Laplace probability distribution. Distributed decoder output M t , aggregate the historical map queries in the streaming memory pool with M t , and then enter the map decoder. The streaming memory pool stores historical feature queries with high confidence. Use the streaming query as the medium for historical feature transmission to transmit information frame by frame, avoiding interaction with all frame images, reducing the computational cost while maintaining high performance. The streaming memory pool is end-to-end and contains all historical queries for scene understanding and joint probabilistic motion planning; In the cross-attention mechanism, the map query interacts with the current frame and historical frames: M′ = MHCA(M t + M t-k , M t-k , M t-k ) Among them, MHCA (M t +M t-k ,M t-k ,M t-k ) represents the cross-attention mechanism, and uses M t +M t-k , M t-k and M t-k as the query, key, and value respectively. M' is the updated query, M t is the updated map query in the distributed decoder, and M t-k is the historical map query, which is composed of map queries with high confidence in the previous frames. Then, the interaction within the current frame is performed in the self-attention mechanism: M″ = MHSA(M′, M′, M′) where MHSA(M′, M′, M′) represents the self-attention mechanism, using M′, M′, and M′ as the query, key, and value respectively, and M″ is the updated query. In the perspective aggregation mechanism, the current scene feature T t is decoded: M = MHPA(M″, T t , T t ) Among them, MHPA(M″,T t ,T t ) represents the perspective aggregation mechanism, using M″, T t and T t as the query, key, and value respectively, and M is the updated map query; After the interaction, the positions and categories of the static map elements are regressed and recognized respectively. To enhance the generalization of the static uncertainty method, on the basis of the two general output structures of point regression and classification head, the regression probability and category probability are added to enable traffic participants to plan their driving following a more reasonable static probability map element.
5. The method for constructing an autonomous driving model according to claim 4, wherein, For the regression probability, the regression head adopts an MLP architecture. For each map element, the regression head generates a two-dimensional vector to represent the point coordinates (x, y) from a bird's-eye view. To convert it into a probability model with uncertainty, a regression head that can output uncertainty parameters related to the predicted points is used; the Laplace distribution is used to model each vertex p = (p1, p2) of the map element. Correspondingly, a map element E with P vertices is expressed as where a represents the a-th vertex, and its joint probability density distribution is where and are the location parameter and scale parameter of the Laplace distribution of the b-th dimension of the a-th vertex of the map element E; Regarding the classification probability, the classification head outputs a class confidence score for each regression vertex. These scores provide a classification distribution in the form of a probability distribution, and this distribution contains sufficient semantic information. Therefore, the probability distribution of these classification scores can be directly passed to the joint probability motion planning without additional processing steps. To encode static uncertainty into joint probabilistic motion planning to reduce such uncertainty errors in planning, for the Laplace distribution of the a-th vertex of the map element E, the map point position parameter λ a , the uncertainty scale parameter s a , and the category parameter l a are concatenated and then encoded using a multi-layer perceptron MLP: Obtained unc p Integrates the vertex features of the probability map; At this time, the updated map query M has static uncertainty. On the one hand, it is used as prior information for planning, and on the other hand, it is used to update the streaming memory pool for future use.
6. The method for constructing an autonomous driving model according to claim 4, wherein The implementation of the object detection in S3.1 includes: The object detection branch is similar to the online map branch, and it is necessary to combine and interact the A obtained by the distributed decoder t with the historical detection queries in the streaming memory pool. Since the pose coordinates of all traffic participants will change before and after movement, the MLN is used to perform pose coordinate transformation on the historical detection queries and A t to update the learnable initial agent query A0 from the scene feature query T t so as to obtain the updated agent query A: A′ = MHCA(MLN(A t + A t-k ), MLN(A t-k ), MLN(A t-k )) A″ = MHSA(A′, A′, A′) A = MHPA(A″, T t , T t ) Finally, a 3D object detection head is used to decode the relevant information of each agent.
7. The method for constructing an autonomous driving model according to claim 1, wherein The joint probability motion planning module in S4 models the agent motion prediction and the ego-vehicle trajectory planning simultaneously as a joint probability motion planning module, and represents the probability distribution with a planning vocabulary and a probability field; First, represent the planned actions as a probability distribution of driving samples, and sample the corresponding actions from the distribution at each time step to control the vehicle. The planned actions are high-dimensional and continuous in time and space; discretize the planned actions into a planned vocabulary cluster and sample N representative planned vocabulary words from it. Represent each action in the planned vocabulary as a trajectory sequence e = (x1, y1, x2, y2,..., x T , y T ), where each trajectory corresponds to a future timestamp; since the actions are continuous on the time axis, the probability corresponding to the action at each moment is continuous; Then, using probability field modeling, the continuous mapping from the planned action to the probability distribution is used to encode the environmental information as Q env , and the planned action is encoded as the initial ego-vehicle motion query After interaction, the final ego-vehicle motion planning query is obtained ego-vehicle query and the environmental information Q env Perform high-order interaction based on Transformer: Among them, the trajectory e is the coordinate value (track sequence) e = (x1, y1, x2, y2,..., x T , y T ), Q hyb is a hybrid instance motion query composed of an agent motion query and a host vehicle motion query, is a historical hybrid instance motion query, and the environmental information Q env includes the hybrid instance motion query Q composed of the host vehicle motion query and the agent motion query hyb , the memory pool hybrid instance motion query detection query A, and the map query M; Input Q for joint probability motion planning hyb There are two sources; the known initial ego - vehicle motion query Using historical detection queries, current detection queries, and current map queries as the initial agent motion query The above two queries serve as the first source; the learnable agent motion query obtained through the distributed decoder and the learnable ego - vehicle motion query These two queries serve as the second source for planning; finally, the queries from the two sources are combined to obtain the input for motion planning; Initial self-vehicle motion query from the first source Expressed as: Among them, (x T , y T ) represents the coordinate position loc, T is the coordinate index, and En() is the position encoding function. En() maps each coordinate loc to a high-dimensional embedding space and is applied to each coordinate value of the trajectory e respectively. En(loc) = concat[ω(loc, 0), ω(loc, 1),..., ω(loc, J - 1)] Inspired by the sine position encoding of the Transformer mechanism, the position encoding function is defined as follows: ω(loc,j) = concat[cos(loc / 10000 2πj / J ,sin(loc / 10000 2πj / J ))] where j represents the encoding dimension index, J is the encoding dimension, ω is the sine-cosine position encoding function that converts the trajectory point coordinates into high-dimensional features, and concat[·] represents concatenation. These functions are used to map the continuous input coordinates to a higher-dimensional space to better approximate the higher-frequency field function, enabling the model to perceive the relative positions of the trajectory points and avoid learning the biases of absolute positions; Combine the current scene understanding information with the prior information in the streaming memory pool to obtain the initialized intelligent agent motion query of the first source Among them, MHCA represents the cross-attention mechanism, and MLP represents the multi-layer perceptron. represents the historical position information of the agent, A m represents the historical detection query, A represents the current detection query, and M represents the current map query. represents the agent learning query embedding; Aggregate the first source and the second source to obtain the agent motion query and the ego vehicle motion query Aggregate the agent motion query and the ego vehicle motion query to obtain the current hybrid instance motion query Q hyb : To fully consider the interactions between the agent and the ego-vehicle, and between agents, an interaction is performed within the current hybrid instance motion query based on the self-attention mechanism: Q hyb = MHSA(Q hyb , Q hyb , Q hyb ) Among them, MHSA(Q hyb ,Q hyb ,Q hyb ) is multi-head self-attention, using Q hyb , Q hyb and Q hyb as query, key, and value respectively; To fully consider the ego-vehicle historical information and the agent historical information, an interaction is performed between the current hybrid instance motion query and the historical hybrid instance motion query: Among them, is cross-attention, using Q hyb , and as query, key, and value respectively; To capture the current agent position and attributes, an interaction is performed between the current hybrid instance motion query and the current detection query: Q hyb = MHCA(Q hyb , A, A) To achieve accurate prediction and planning, both the agent and the ego-vehicle need to consider the high-level semantic information containing the static uncertainty map, and an interaction is performed between the current hybrid instance motion query and the online map query: Q hyb = MHCA(Q hyb , M, M) Among them, MHCA(Q hyb , M, M) is multi-head cross attention, using Q hyb , M, and M as query, key, and value respectively; Q hyb is the updated current mixed instance motion query, which contains the current agent motion query Q a and the ego vehicle motion query Q e ; The learned hybrid instance motion query fuses the high-order ego-vehicle-environment-temporal interaction relationships, including the uncertainty static representation and the uncertainty dynamic representation. Finally, the corresponding probability planning result is calculated through MLP: In the multi-modal trajectories at different times, the maximum probability trajectory e* = argmax p(e) is adopted as the current trajectory to reduce the static uncertainty and the dynamic uncertainty; meanwhile, the updated hybrid instance motion query is used to update the streaming memory pool to enter the planning of the next frame.
8. The method for constructing an autonomous driving model according to claim 1, wherein It also includes a designed loss function, and the loss function includes three task loss functions, namely detection loss L det , map loss L map , and joint probability motion planning loss L motion_plan , as follows: L = λ1L det + λ2L map + λ3L motion_plan where λ1, λ2, and λ3 are the corresponding loss weights; Detection loss L det is the classification loss L FocalLoss and the regression loss L L1 which is a linear combination of: L det_cls = L FocalLoss , L det_reg = L L1 L det = L det_cls + L det_reg The map loss consists of the map element classification loss L Focal and the negative log-likelihood loss L NLL for map instance regression: L map_cls = L Focal , L map_reg = L NLL L map = L map_cls + L map_reg In joint probability motion planning, the ground truth trajectory is added to the planning vocabulary as a positive sample, and the remaining trajectories are negative samples. The KL divergence is used to calculate the distribution loss between the predicted distribution and the ground truth distribution. In addition, the ego-vehicle-agent collision constraint L is added. ea 、the ego-vehicle-road boundary crossing constraint L eb 、the ego-vehicle-lane direction constraint L el are as follows:
9. An end-to-end autonomous driving model that combines distributed queries and fuses the uncertainty of spatio-temporal information, characterized in that The autonomous driving model is obtained by the construction method of claim 1.
10. A vehicle-mounted device, characterized in that, The device deploys the model of claim 9.
Citation Information
Cited By
Training method and device of trajectory prediction model and medium
CN120656003A