End-to-end automatic driving system based on multi-modal fusion
By using a multimodal fusion-based end-to-end autonomous driving system, and leveraging an improved feature extraction and fusion network, combined with an attention mechanism and a deep Q-network, the problems of error accumulation and low learning efficiency in traditional systems are solved, achieving higher precision in environmental perception and decision-making capabilities.
Patent Information
- Application Number
- CN202511070082.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-18
Smart Images

Figure CN120963755A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of automatic driving, and particularly relates to an end-to-end automatic driving system based on multi-modal fusion. BACKGROUND
[0002] With the development of artificial intelligence technology, end-to-end automatic driving systems have been widely researched and applied in the global range. Traditional automatic driving systems have multiple independent sub-modules between sensor input and actuator output, such as environment perception, path planning, motion control, etc. However, this kind of modular solution has the problem of cumulative error in the process of executing tasks, which affects driving safety. In order to reduce the cumulative error in the execution process, researchers have proposed the concept of end-to-end automatic driving, which maps the original sensor data to the future driving trajectory or control instruction through a neural network model. Compared with the traditional method, the end-to-end automatic driving scheme not only simplifies the decision-making process, but also reduces the problems caused by error transmission between modules.
[0003] Early end-to-end automatic driving uses a monocular camera, but there are great limitations in automatic driving decision-making from a single modality to obtain environmental image information. In recent years, people have installed multiple complementary sensors on the car, such as cameras, LiDARs and radars, etc. In order to realize the fusion of different sensor information, existing researches mostly project the features of one modality onto another modality to realize modality enhancement, but this way may cause the loss of information of the projected modality. Some researches also try to associate the features of different sensors through attention mechanism, but this attention learning efficiency is low through feature similarity in different views.
[0004] In addition, the existing end-to-end automatic driving model mostly adopts a complex feature extraction and fusion network, and the behavior planning network structure is only composed of several simple multi-layer perceptrons (MLP) and gated recurrent units (GRU). This imbalance between resources and tasks limits the learning efficiency of the model and affects the overall learning ability of the model. Even if the feature extraction and fusion network can provide high-quality feature representation, due to the limited ability of the behavior planning network, the model may not be able to effectively learn the mapping relationship from input to output. SUMMARY
[0005] The present application relates to the technical field of automatic driving, and particularly relates to an end-to-end automatic driving system based on multi-modal fusion.
[0006] The present application relates to the technical field of automatic driving, and particularly relates to an end-to-end automatic driving system based on multi-modal fusion.
[0007] An end-to-end autonomous driving system based on multimodal fusion includes:
[0008] The environmental perception module includes various types of sensors for collecting environmental information in different modalities;
[0009] The feature extraction module is used to extract features from the collected environmental information of different modalities and transform it into a bird's-eye view space to obtain multimodal BEV features;
[0010] The feature fusion module is used to fuse multimodal BEV features to obtain fused features;
[0011] The behavior planning module is used to use a trained behavior planning network to output the predicted future waypoints and control commands for the target vehicle based on the fusion features.
[0012] The output fusion module is used to perform trajectory prediction and multi-step control prediction on the future waypoints and control commands output by the behavior planning module, and select the optimal driving behavior based on the context.
[0013] Furthermore, the various types of sensors include cameras, LiDAR devices, and Radar devices.
[0014] Furthermore, the feature extraction module employs an improved PointNet++ network to extract the 3D point cloud features acquired by the LiDAR device, specifically including:
[0015] Before the 3D point cloud features acquired by the LiDAR device are input into the improved PointNet++ network, noise and outliers in the 3D point cloud features are removed by statistical filtering. The statistical filtering method includes: for each point in the 3D point cloud features, calculating the local mean and standard deviation in the neighborhood of each point, thereby calculating the Z score of each point in each principal component direction. If the Z score exceeds the preset Z threshold, it is judged as an outlier and removed.
[0016] The noise- and outlier-removed 3D point cloud features are input into an improved PointNet++ network, in which the 1D convolutions of the feature extraction layer are replaced with separable convolutions in the STN3D and STNkD modules of the PointNet++ network.
[0017] The improved PointNet++ network performs convolution operations independently on each channel of the input 3D point cloud features to generate multiple independent feature maps; it uses 1x1 convolution kernels to combine the independent feature maps of each channel to achieve cross-channel feature input, and after global max pooling, it obtains the extracted global feature map of the LiDAR device point cloud.
[0018] Furthermore, the feature extraction module uses a Vision Mamba network to extract image features from multi-view image data captured by the camera, specifically including:
[0019] The multi-view image data captured by the camera is cropped, scaled, and normalized to ensure that the input data is of uniform specifications.
[0020] The multi-view image data captured by the camera is divided into multiple patches of the same size. Each patch is flattened and mapped to a D-dimensional vector. At the same time, position encoding is added to each patch to obtain a patch sequence.
[0021] The processed patch sequence is input into the Vision Mamba encoder, which consists of multiple identical stacked layers. Each layer contains a bidirectional state space model and a feedforward network. The bidirectional state space model is used to process the sequence data bidirectionally and capture the global dependencies between patches; the feedforward network is used to further extract features.
[0022] The Vision Mamba encoder outputs the feature representations of each viewpoint, which are then stitched together to obtain the final multi-view fused image features.
[0023] Furthermore, the feature fusion module's fusion process includes a preliminary fusion step and an adaptive fusion step, the preliminary fusion step including:
[0024] The features extracted by the feature extraction module for different modes are mapped onto the BEV coordinate system centered on the vehicle to obtain the BEV features for different modes.
[0025] An initial BEV query is generated based on the target vehicle's location code;
[0026] Multi-head attention is used to fuse BEV features from different modalities to obtain fused features from different modalities. These features are then concatenated and further fused through convolutional layers to obtain the final fused BEV features.
[0027] The adaptive fusion step includes:
[0028] The state is defined based on the fusion features of each modality and the BEV fusion features at the current moment; the action is defined based on the fusion weights of the fusion features of each modality; and the reward function is defined based on the consistency between the fusion features of each modality and the BEV query, as well as the stability of the fusion features of adjacent frames.
[0029] Based on the defined state, action, and reward functions, a deep Q-network is used for adaptive fusion to dynamically adjust the fusion weights of multimodal sensor features;
[0030] During the training of the deep Q-network, an ε-greedy strategy is used in the initial training phase to randomly select actions with ε probability in the state and select actions that maximize the Q value with 1-ε probability. After the action is executed, the reward and the state at the next moment are observed, and experience is generated and stored in the replay buffer. Batch experience is randomly sampled to train the deep Q-network. The loss function in the training process of the deep Q-network includes multi-sensor fusion network loss, target detection task loss and total loss function.
[0031] Furthermore, the processing procedure of the behavior planning module includes:
[0032] A query vector is constructed, and then position encoding fusion is performed. The query vector includes trajectory point queries and control command queries. The trajectory point queries use local target points as initial values, representing possible future path points of the vehicle; the control command queries are initialized to a zero vector or preset basic control values. Learnable position codes are added to the query vector to obtain the fused query vector. The corresponding calculation expression is:
[0033]
[0034] In the formula, This is the fused query vector at the initial time step. For querying trajectory points at the initial moment, For initial control command query, PE motion For learnable location encoding;
[0035] By capturing the intrinsic relationship between trajectory point queries and control command queries through a self-attention mechanism, the fused query vector is updated using self-attention. The corresponding calculation expression is as follows:
[0036]
[0037] In the formula, Let be the updated query vector at time n. For layer normalization, It is a multilayer perceptron, and Self-Attenuation is a self-attention computation. This is the query vector at time n after fusion;
[0038] Using a cross-attention mechanism, the updated query vector interacts with the fused features to dynamically update the query vector. The corresponding calculation expression is:
[0039]
[0040] In the formula, This is the updated query vector at time n+1. To fuse features, Cross-Atten extracts scene information relevant to the current decision by calculating the attention weights of the query and BEV features;
[0041] The query vector is iteratively updated to obtain the final query vector, and then connected to the prediction head to predict multiple future waypoints and control commands.
[0042] Furthermore, the output fusion module includes a trajectory prediction branch and a multi-step control prediction branch, and adopts a context-based fusion strategy to determine the final prediction result;
[0043] The trajectory prediction branch uses a diffusion model to generate future trajectories, and then uses an MPC controller for lateral control and a PID controller for longitudinal control.
[0044] The multi-step control prediction branch adopts a trajectory-guided attention mechanism, obtains information from the trajectory prediction branch, predicts the control actions of the next multiple steps, and directly predicts the vehicle's steering value, throttle value, and braking value.
[0045] The context-based fusion strategy fuses the outputs of the trajectory prediction branch and the multi-step control prediction branch by weighted averaging, and adjusts the fusion weights according to the vehicle's driving state to obtain the final control signal.
[0046] Furthermore, the multi-step control prediction branch uses a GRU model to output the updated hidden state based on the current features and the predicted control action.
[0047] Hidden states of prediction branches using multi-step control Hidden states of trajectory prediction branches The attention map is calculated using MLP, and then...
[0048] Image features F extracted by the attention map aggregation image feature extraction module MV Hidden states of multi-step control prediction branches The resulting concatenated feature is an enhanced feature used for predicting future control actions in the multi-step control prediction branch. The calculation expression for the enhanced feature is as follows:
[0049]
[0050] In the formula, To enhance the features, t represents time.
[0051] Furthermore, the calculation rule for the final control signal obtained by the context fusion-based strategy is as follows:
[0052] If the target vehicle is in a trajectory-specific scenario, then the expression for the final control signal α is:
[0053] a=α·a ctl +(1-α)·a traj
[0054] In the formula, α is the fusion weight, a ctl To control the output action of the prediction branch in a multi-step process, a traj The output action for the trajectory prediction branch;
[0055] If the target vehicle is in a dedicated control scenario, then the expression for the final control signal α is:
[0056] a=α·a traj +(1-α)·a ctl .
[0057] Furthermore, the system also includes a central hardware trigger, a time stamp module, a data buffer module, and a data alignment module;
[0058] The central hardware trigger is used to send synchronization signals to each sensor in the environmental perception module, thereby triggering the operation of each sensor in a unified manner.
[0059] The time stamping module is connected to the environment sensing module and is used to mark the environmental information data of each modality with timestamps.
[0060] The data buffer module is connected to the time stamp module and is used to store environmental information data with timestamps.
[0061] The data alignment module is connected to the data buffer module and is used to find and align the environmental information data of each modality in the data buffer module according to the timestamp.
[0062] Compared with the prior art, the present invention has the following advantages:
[0063] (1) This paper proposes to extract features from environmental information data of multiple modalities and then use bird’s-eye view (BEV) representation for fusion. By establishing a unified coordinate system and scale for all sensors, it makes data association between different sensors easier, greatly reducing the complexity of system design and promoting integration with other sensors. In the fusion process, two steps are adopted: preliminary fusion and adaptive fusion. On the one hand, preliminary fusion is carried out through a multi-head attention mechanism, and on the other hand, adaptive fusion is carried out based on a deep Q network to dynamically adjust the fusion weights of multimodal sensor features, which can quickly obtain fused features that achieve feature consistency and stability.
[0064] Furthermore, the behavior planning module integrates predicted waypoints and control commands, enhancing the model's generalization ability in complex scenarios.
[0065] (2) In the behavior planning module of the present invention, on the one hand, the trajectory point query and control command query are fused by joint motion query and a learnable position code is added; on the other hand, the self-attention mechanism is used to capture the intrinsic relationship between the trajectory point query and the control command query, and the cross-attention mechanism is used to calculate the attention weight of the query and the fused BEV features, extract the scene information related to the current decision, and gradually fuse the global scene information during the iteration process, which effectively improves the accuracy of the decision.
[0066] (3) The output fusion module of the present invention is equipped with a trajectory prediction branch and a multi-step control prediction branch. The trajectory prediction branch outputs the predicted future trajectory waypoints and guides the multi-step control branch to predict the control signals for the next multiple steps. Traditional control prediction methods can only predict the control commands for a single step in the future. The present invention introduces a predicted trajectory to guide the control prediction head to predict the control signals for the current and the next K steps. During the prediction process of the multi-step control prediction branch, the hidden states of the multi-step control prediction branch and the trajectory prediction branch are used to calculate the attention map through MLP. The image features extracted by the image feature extraction module are aggregated through the attention map and then spliced with the hidden states of the control branch to generate enhanced features, thereby achieving accurate prediction of the control signals for the next multiple steps.
[0067] (4) The output fusion module of the present invention also proposes to determine whether the current scene is more suitable for relying on trajectory planning or direct control based on the vehicle driving status, and dynamically adjust the weights of the two to avoid the limitations of a single branch. Attached Figure Description
[0068] Figure 1 This is a schematic diagram of the structure of an end-to-end autonomous driving system based on multimodal fusion provided in an embodiment of the present invention;
[0069] Figure 2 This is a flowchart illustrating the overall workflow of an end-to-end autonomous driving system based on multimodal fusion, as provided in an embodiment of the present invention.
[0070] Figure 3 This is a flowchart illustrating the behavior planning module and output fusion module of an end-to-end autonomous driving system based on multimodal fusion, as provided in an embodiment of the present invention. Detailed Implementation
[0071] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0072] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0073] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0074] Example 1
[0075] This embodiment provides an end-to-end autonomous driving system based on multimodal fusion, relating to the field of autonomous driving technology. The end-to-end autonomous driving system based on multimodal fusion provided in this embodiment can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the end-to-end autonomous driving system based on multimodal fusion, but is not limited to the above forms.
[0076] like Figure 1 and Figure 2 As shown, it specifically includes:
[0077] The environmental perception module 101 includes various types of sensors for collecting environmental information in different modalities; the environmental information includes static and dynamic information. As an input to end-to-end autonomous driving, the environmental perception module 101 is responsible for perceiving the environmental information around the vehicle, thereby achieving a comprehensive understanding of the surrounding environment.
[0078] The feature extraction module 102 is used to extract features from the collected environmental information of different modalities and transform it into the bird's-eye view space to obtain multimodal BEV features.
[0079] The feature fusion module 103 is used to fuse the extracted features from different modalities to obtain fused features;
[0080] The process includes preliminary fusion and adaptive fusion. Preliminary fusion generates an initial BEV query based on position encoding and uses a multi-head attention mechanism to interact with the BEV features of each modality to obtain fused modal features. Then, the fused features of each modality are concatenated and passed through a convolutional layer to obtain the final BEV fused features. Adaptive fusion adopts a reinforcement learning method based on a DQN network to dynamically adjust the fusion weights of multimodal sensor features. The fused BEV features of each modality and the current fused features are used as the state, and the dynamic adjustment weights of each modal feature are used as the action. A reward function based on feature consistency and stability is designed.
[0081] The behavior planning module 104 is used to use a trained behavior planning network to output the predicted results of the future waypoints and control commands of the target vehicle based on the fusion features.
[0082] The output fusion module 105 is used to perform trajectory prediction and multi-step control prediction on the future waypoints and control commands output by the behavior planning module, and select the optimal driving behavior based on the context; that is, it includes trajectory prediction branch and multi-step control prediction branch, and uses a context-based fusion strategy to determine which result is more dependent.
[0083] In this embodiment,
[0084] The environmental perception module 101 acquires vehicle and surrounding data through cameras, LiDAR, Radar, GPS, inertial sensors, etc.
[0085] Specifically, the camera is used to collect environmental image data;
[0086] LiDAR is used to acquire 3D point cloud data of the environment;
[0087] Radar is used to collect the motion state of dynamic objects in the environment;
[0088] GPS is used to collect the absolute location information of vehicles;
[0089] Inertial sensors are used to collect dynamic status information of vehicles.
[0090] Preferably, the system also includes a central hardware trigger, a time stamp module, a data buffer module, and a data alignment module;
[0091] The central hardware trigger is used to send synchronization signals to each sensor in the environmental perception module, thereby triggering the operation of each sensor in a unified manner.
[0092] The time stamping module is connected to the environment perception module and is used to add timestamps to the environmental information data of each modality;
[0093] The data buffer module is connected to the timestamp module and is used to store environmental information data with timestamps.
[0094] The data alignment module is connected to the data buffer module and is used to find and align the environmental information data of each modality within the data buffer module based on the timestamp.
[0095] The feature extraction module 102 extracts the 3D point cloud features of LiDAR through the improved PointNet++, extracts the image features of multi-view image data through the VisionMamba network, and extracts the point cloud features of Radar through PointPillars.
[0096] An improved PointNet++ network is used to extract features from 3D point clouds acquired by LiDAR devices, specifically including:
[0097] Before inputting the 3D point cloud features acquired by the LiDAR device into the improved PointNet++ network, noise and outliers in the 3D point cloud features are removed by statistical filtering methods. The statistical filtering methods include: for each point in the 3D point cloud features, calculating the local mean and standard deviation in the neighborhood of each point, thereby calculating the Z score of each point in each principal component direction. If the Z score exceeds the preset Z threshold, it is judged as an outlier and removed.
[0098] The noise- and outlier-removed 3D point cloud features are input into the improved PointNet++ network. In the STN3D and STNkD modules of the PointNet++ network, the 1D convolutions of the feature extraction layer are replaced with separable convolutions.
[0099] The improved PointNet++ network performs convolution operations independently on each channel of the input 3D point cloud features to generate multiple independent feature maps. It uses 1x1 convolution kernels to combine the independent feature maps of each channel to achieve cross-channel feature input. After global max pooling, it obtains the extracted global features of the LiDAR device's point cloud.
[0100] In this embodiment, an improved PointNet++ is used as the backbone network for LiDAR point cloud extraction, with improvements mainly made in the preprocessing stage and the network structure.
[0101] Improvements in the preprocessing stage: Before inputting LiDAR point cloud data into the PointNet++ network, statistical filtering methods are used to remove noise and outliers, thereby improving data quality.
[0102] For each point, calculate the local mean μ in its neighborhood. i and standard deviation σ i The formula is as follows:
[0103]
[0104] Where N i Representative point x i The set of neighboring points; k represents the number of neighboring points.
[0105] Then calculate the Z-score for each point along each principal component direction, using the following formula:
[0106]
[0107] If the Z-score exceeds the preset threshold (2-3 times the standard deviation), it is identified as an outlier and removed.
[0108] The network structure is improved in PointNet++'s STN3D and STNkD modules by replacing the traditional 1D convolutions in the feature extraction layer with separable convolutions, reducing computational complexity.
[0109] Perform convolution operations independently on each channel of the input. If the input has D channels, then D independent feature maps are generated, as shown in the following formula:
[0110]
[0111] in For the output feature of the d-th channel at position t, W d,Γ These are the corresponding convolution kernel weights.
[0112] Global max pooling is used to obtain the global feature F of the LiDAR point cloud. LiDAR .
[0113] Feature extraction module 102 uses a Vision Mamba network to extract image features from multi-view image data captured by the camera, specifically including:
[0114] The multi-view image data captured by the camera is cropped, scaled, and normalized to ensure that the input data is of uniform specifications.
[0115] The multi-view image data captured by the camera is divided into multiple patches of the same size. Each patch is flattened and mapped to a D-dimensional vector. At the same time, position encoding is added to each patch to obtain a patch sequence.
[0116] The processed patch sequence is fed into the Vision Mamba encoder, which consists of multiple stacked layers. Each layer contains a bidirectional state space model and a feedforward network. The bidirectional state space model is used to process the sequence data bidirectionally and capture the global dependencies between patches; the feedforward network is used to further extract features.
[0117] The Vision Mamba encoder outputs the feature representations of each viewpoint, which are then stitched together to obtain the final multi-view fused image features.
[0118] In this embodiment, the Vision Mamba network first performs cropping, scaling, and normalization processing on images from different perspectives to ensure that the input data specifications are uniform.
[0119] Next, the image is divided into patches of equal size. Each patch is flattened and mapped to a D-dimensional vector, preparing it for subsequent processing. For example, an input image of size H×W×C will be divided into (H×W) / P patches. 2 Each patch is then normalized, adjusting its range to [0,1] or [-1,1].
[0120] Simultaneously, a positional encoding (MVPE) is added to each patch, enabling the model to understand the spatial location information of the patch in the image and enhancing its feature representation ability. The processed patch sequence is then fed into the Vision Mamba encoder, which consists of multiple stacked identical layers, each containing a bidirectional state-space model (BSSM) and a feedforward network (FFN).
[0121] BSSM is responsible for bidirectional processing of sequence data and capturing global dependencies between patches;
[0122] FFN further extracts features. After this series of processes, feature representations for each viewpoint are obtained.
[0123] Finally, the features from each perspective are stitched together to obtain the final multi-view fused image feature F. MV This enables the comprehensive utilization of image information from different perspectives, providing image feature data for subsequent fusion and target detection in bird's-eye view (BEV) space.
[0124] PointPillars extracts point cloud features from Radar to filter the point cloud range. Based on the detection range of Radar, invalid points are filtered out. The three-dimensional spatial range is set as x∈[0,51.2]m, y∈[-25.6,25.6]m, z∈[-3,2]m, focusing on the effective detection area.
[0125] Then, the selected voxels (e.g., 0.16m×0.16m×0.24m) are treated as pillars, reducing the amount of data and structuring the features.
[0126] For each radar point within a voxel, features such as relative coordinates, original coordinates, and reflection intensity are calculated, and the point features are encoded into high-dimensional vectors using a multilayer perceptron (MLP).
[0127] The features of all points within each voxel are fused to generate a feature vector for a single pillar, compressing the point cloud information while retaining key features. This yields the point cloud features F of the Radar. radar .
[0128] The feature fusion module 103's fusion process includes a preliminary fusion step and an adaptive fusion step. The preliminary fusion step includes:
[0129] The features extracted by the feature extraction module for different modes are mapped onto the BEV coordinate system centered on the vehicle to obtain the BEV features for different modes.
[0130] An initial BEV query is generated based on the target vehicle's location code;
[0131] Multi-head attention is used to fuse BEV features from different modalities to obtain fused features from different modalities. These features are then concatenated and further fused through convolutional layers to obtain the final fused BEV features.
[0132] The adaptive fusion steps include:
[0133] The state is defined based on the fusion features of each modality and the BEV fusion features at the current moment; the action is defined based on the fusion weights of the fusion features of each modality; and the reward function is defined based on the consistency between the fusion features of each modality and the BEV query, as well as the stability of the fusion features of adjacent frames.
[0134] Based on the defined state, action, and reward functions, a deep Q-network is used for adaptive fusion to dynamically adjust the fusion weights of multimodal sensor features;
[0135] During the training of the deep Q-network, an ε-greedy strategy is used in the initial training phase to randomly select actions with ε probability in the state and select the action that maximizes the Q value with 1-ε probability. After the action is executed, the reward and the state at the next time step are observed, and the experience is generated and stored in the replay buffer. Batch experience is randomly sampled to train the deep Q-network. The loss function in the training process of the deep Q-network includes the multi-sensor fusion network loss, the object detection task loss and the total loss function.
[0136] Specifically, in this implementation, the feature fusion module is connected to the feature extraction module 102 and the behavior planning module 104 respectively. The feature fusion module 103 includes two steps: preliminary fusion and adaptive fusion. First, the BEV features output after feature transformation are initially fused using Transformer. Then, an initial BEV query Q is generated based on the position encoding. BEV .
[0137] Then, multi-head attention (MHSA) fusion is employed. Taking image features as an example, through... Generate queries, keys, and values, calculate self-attention scores, and fuse features to obtain...
[0138] Similarly, by processing LiDAR and Radar features, we obtain and Finally, the fused features of the three modalities are concatenated into F. concat-BEV The final fused features are obtained by further fusing them through convolutional layers.
[0139] Adaptive fusion is performed using a deep Q-network (DQN) based system, defining the state. Includes modal fusion features and BEV fusion features at the current moment, defining actions. This indicates the adjustment of the fusion weights for each modality feature.
[0140] Define reward function R t =σ1·Consistency t +σ2·Stability t Consistency t Stability measures the consistency between each modal feature and the BEV query. t To measure the stability of the fused features of adjacent frames, negative distance calculation is used to ensure that the smaller the distance, the higher the reward.
[0141] During the training of DQN, the weights of the Q-network and the target network are first initialized, and the initial state s0 is set. Then, an ε-greedy policy is used in policy selection in state s0. t Select action a t We randomly select an action with probability ε, and select the action that maximizes the Q value with probability 1-ε.
[0142] Finally, experience replay and updates are performed. After executing an action, the reward and the next state are observed, and the experience is stored in the replay buffer. Batch experiences are randomly sampled to train the network, and the Bellman equation Q(s) is used. t ,a t )←Q(s t ,a t )+α[Rt +γmax a′ Q(s t+1 ,a′)-Q(s t ,a t Update the Q value and periodically synchronize the target network weights.
[0143] For the loss function, we designed the multi-sensor fusion network loss, the target detection task loss, and the total loss function respectively.
[0144] The consistency loss of a multi-sensor fusion network is This is to ensure that the fused features retain important information from each modality. The stability loss is... To ensure the stability of the fused features over time, the fusion loss is L. fusion =λ1L consistency +λ2L stability To balance consistency and stability.
[0145] The classification loss for the object detection task is the cross-entropy loss. Calculate the difference between the predicted category and the true category.
[0146] The positioning loss uses smoothed L1 loss. Measure the positional error between the predicted bounding box and the ground truth bounding box. Detection loss L dct =λ3L cls +λ4L loc It is used to comprehensively optimize classification and positioning.
[0147] The total loss function is expressed as L total =L fusion +L dct It is used to optimize fusion networks and detection tasks.
[0148] The processing steps of the behavior planning module 105 include:
[0149] A query vector is constructed, and then position codes are fused. The query vector includes trajectory point queries and control command queries. The trajectory point queries use local target points as initial values, representing possible future path points of the vehicle. The control command queries are initialized to a zero vector or preset basic control values. Learnable position codes are added to the query vector to obtain the fused query vector.
[0150] By capturing the intrinsic relationship between trajectory point queries and control command queries through a self-attention mechanism, the fused query vector is updated using self-attention.
[0151] The query vector is dynamically updated by interacting with the fused features using a cross-attention mechanism.
[0152] The query vector is iteratively updated to obtain the final query vector, and then connected to the prediction head to predict multiple future waypoints and control commands.
[0153] In this embodiment, as Figure 3 As shown, the behavior planning module 105 is connected to the feature fusion module 104 and the output fusion module 106. The behavior planning module is equipped with a joint motion query and attention-driven motion planning network.
[0154] The joint motion query includes trajectory queries and control command queries, used to extract the scene context required for decision-making from BEV features.
[0155] Attention-driven motion planning networks dynamically optimize the interaction between queries and BEV features through self-attention and cross-attention mechanisms.
[0156] The joint motion query first constructs a query vector, including trajectory point queries and control command queries, and then performs position encoding fusion.
[0157] Trajectory point query is represented as The local target point is used as the initial value to represent the possible future path points of the vehicle.
[0158] Control command query is represented as Initialize to a zero vector or a preset base control value.
[0159] Location coding fusion will enable learnable location coding PE motion Add it to the query vector, as shown in the following formula:
[0160]
[0161] This encoding helps the model distinguish the temporal and spatial relationships between different queries.
[0162] Attention-driven query optimization involves using self-attention to establish relationships between queries, then fusing BEV scenario context through cross-attention, and finally optimizing through multi-level iterations.
[0163] The intrinsic relationship between trajectory point queries and control command queries is captured through a self-attention mechanism, as shown in the following formula:
[0164]
[0165] in, For layer normalization, It is a multilayer perceptron that maintains information integrity through residual connections.
[0166] Using a cross-attention mechanism, the updated query vector and the fused BEV vector are... Interactive, dynamically updated query, formula as follows:
[0167]
[0168] Cross-Atten extracts scene information (such as obstacle location and road structure) relevant to the current decision by calculating the attention weights between the query and BEV features.
[0169] Through iterative processing of multi-layered attention modules, query vectors gradually integrate global scene information, improving the accuracy of decision-making.
[0170] The output fusion module 106 includes a trajectory prediction branch and a multi-step control prediction branch, and adopts a context-based fusion strategy to determine the final prediction result.
[0171] The trajectory prediction branch uses a diffusion model to generate future trajectories, then uses an MPC controller for lateral control and a PID controller for longitudinal control.
[0172] The multi-step control prediction branch adopts a trajectory-guided attention mechanism, obtains information from the trajectory prediction branch, predicts the control actions of the next multiple steps, and directly predicts the vehicle's steering value, throttle value, and braking value.
[0173] The context-based fusion strategy fuses the outputs of the trajectory prediction branch and the multi-step control prediction branch by weighted averaging, and adjusts the fusion weights according to the vehicle's driving state to obtain the final control signal.
[0174] In this embodiment, the output fusion module 106 is connected to the behavior planning module 105. This module has two branches: a trajectory prediction branch and a multi-step control prediction branch. The trajectory prediction branch outputs the predicted future trajectory waypoints while also guiding the multi-step control branch to predict the control signals for the next multiple steps.
[0175] Two prediction heads are connected after the output of the behavior planning module 105 to predict multiple future waypoints and control commands, respectively. The waypoint prediction head uses a single-layer GRU network, taking four waypoint query embeddings as input, and employs a diffusion model to predict multiple future waypoints. Then, the MPC controller performs lateral control, and the PID controller performs longitudinal control, generating the vehicle's steering value (steer). wp ∈[-1,1], throttle value wp ∈[-1,1] and brake value wp ∈[-1,1].
[0176] Define the future K-step path points generated by the diffusion model as follows:
[0177] Define the control signal that transforms the trajectory using MPC and PID as a traj .
[0178] The loss function for the trajectory planning branch is as follows:
[0179]
[0180] Among them, wp t The path point at the t-th time step predicted by the trajectory planning branch; Represents the actual path point at time step t; The characteristic representation of the trajectory branch at the current time (t=0); The feature representation of the expert model at the current moment.
[0181] Traditional control prediction methods can only predict control commands for a single future step. This invention introduces a prediction trajectory to guide the control prediction head in predicting control signals for the current and future K steps.
[0182]
[0183] Each action includes accelerator, brake, and steering.
[0184] Dynamic interaction is modeled using GRU, with the current feature as the input. With predicting action a t Output the updated hidden state Implicit simulation of the dynamic changes in the environment and vehicles.
[0185] GRU hidden states using trajectory branches and the hidden state of control branches The attention map is calculated using an MLP, as shown in the following formula:
[0186]
[0187] Then, feature fusion is performed, aggregating the image features F extracted by the image feature extraction module through attention maps. MV Enhanced features are generated by concatenating the hidden states of the control branches. The formula is as follows:
[0188]
[0189] Based on the vehicle's driving status, determine whether the current scenario is more suitable for trajectory planning or direct control, and dynamically adjust the weights of both to avoid the limitations of a single branch.
[0190] Let the fusion weight be α∈[0,0.5], and the calculation rule for the final control signal α is as follows:
[0191] If in a trajectory-specific context:
[0192] a=α·a ctl +(1-α)·a traj
[0193] If in a control-specific context:
[0194] a=α·a traj +(1-a)·a ctl
[0195] Here, (1-α) is the weight of the dominant branch, ensuring that the branch more suitable for the current scenario dominates.
[0196] In addition, advanced navigation instructions are provided by using A * The algorithm generates a global trajectory, including straight, left turn, right turn, and lane keeping.
[0197] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. An end-to-end autonomous driving system based on multimodal fusion, characterized in that, include: The environmental perception module includes various types of sensors for collecting environmental information in different modalities; The feature extraction module is used to extract features from the collected environmental information of different modalities and transform it into a bird's-eye view space to obtain multimodal BEV features; The feature fusion module is used to fuse multimodal BEV features to obtain fused features; The behavior planning module is used to use a trained behavior planning network to output the predicted future waypoints and control commands for the target vehicle based on the fusion features. The output fusion module is used to perform trajectory prediction and multi-step control prediction on the future waypoints and control commands output by the behavior planning module, and select the optimal driving behavior based on the context.
2. The end-to-end autonomous driving system based on multimodal fusion according to claim 1, characterized in that, The various types of sensors include cameras, LiDAR devices, and Radar devices.
3. The end-to-end autonomous driving system based on multimodal fusion according to claim 2, characterized in that, The feature extraction module uses an improved PointNet++ network to extract 3D point cloud features acquired by the LiDAR device, specifically including: Before the 3D point cloud features acquired by the LiDAR device are input into the improved PointNet++ network, noise and outliers in the 3D point cloud features are removed by statistical filtering. The statistical filtering method includes: for each point in the 3D point cloud features, calculating the local mean and standard deviation in the neighborhood of each point, thereby calculating the Z score of each point in each principal component direction. If the Z score exceeds the preset Z threshold, it is judged as an outlier and removed. The noise- and outlier-removed 3D point cloud features are input into an improved PointNet++ network, in which the 1D convolutions of the feature extraction layer are replaced with separable convolutions in the STN3D and STNkD modules of the PointNet++ network. The improved PointNet++ network performs convolution operations independently on each channel of the input 3D point cloud features to generate multiple independent feature maps; it uses 1x1 convolution kernels to combine the independent feature maps of each channel to achieve cross-channel feature input, and after global max pooling, it obtains the extracted global feature map of the LiDAR device point cloud.
4. An end-to-end autonomous driving system based on multimodal fusion according to claim 2, characterized in that, The feature extraction module uses a Vision Mamba network to extract image features from multi-view image data captured by the camera, specifically including: The multi-view image data captured by the camera is cropped, scaled, and normalized to ensure that the input data is of uniform specifications. The multi-view image data captured by the camera is divided into multiple patches of the same size. Each patch is flattened and mapped to a D-dimensional vector. At the same time, position encoding is added to each patch to obtain a patch sequence. The processed patch sequence is input into the Vision Mamba encoder, which consists of multiple identical stacked layers. Each layer contains a bidirectional state space model and a feedforward network. The bidirectional state space model is used to process the sequence data bidirectionally and capture the global dependencies between patches; the feedforward network is used to further extract features. The Vision Mamba encoder outputs the feature representations of each viewpoint, which are then stitched together to obtain the final multi-view fused image features.
5. An end-to-end autonomous driving system based on multimodal fusion according to claim 2, characterized in that, The feature fusion module's fusion process includes a preliminary fusion step and an adaptive fusion step. The preliminary fusion step includes: The features extracted by the feature extraction module for different modes are mapped onto the BEV coordinate system centered on the vehicle to obtain the BEV features for different modes. An initial BEV query is generated based on the target vehicle's location code; Multi-head attention is used to fuse BEV features from different modalities to obtain fused features from different modalities. These features are then concatenated and further fused through convolutional layers to obtain the final fused BEV features. The adaptive fusion step includes: The state is defined based on the fusion features of each modality and the BEV fusion features at the current moment; the action is defined based on the fusion weights of the fusion features of each modality; and the reward function is defined based on the consistency between the fusion features of each modality and the BEV query, as well as the stability of the fusion features of adjacent frames. Based on the defined state, action, and reward functions, a deep Q-network is used for adaptive fusion to dynamically adjust the fusion weights of multimodal sensor features; During the training of the deep Q-network, an ε-greedy strategy is used in the initial training phase to randomly select actions with ε probability in the state and select actions that maximize the Q value with 1-ε probability. After the action is executed, the reward and the state at the next moment are observed, and experience is generated and stored in the replay buffer. Batch experience is randomly sampled to train the deep Q-network. The loss function in the training process of the deep Q-network includes multi-sensor fusion network loss, target detection task loss and total loss function.
6. An end-to-end autonomous driving system based on multimodal fusion according to claim 1, characterized in that, The processing steps of the behavior planning module include: A query vector is constructed, and then position encoding fusion is performed. The query vector includes trajectory point queries and control command queries. The trajectory point queries use local target points as initial values, representing possible future path points of the vehicle; the control command queries are initialized to a zero vector or preset basic control values. Learnable position codes are added to the query vector to obtain the fused query vector. The corresponding calculation expression is: In the formula, This is the fused query vector at the initial time step. For querying trajectory points at the initial moment, For initial control command query, PE motion For learnable location encoding; By capturing the intrinsic relationship between trajectory point queries and control command queries through a self-attention mechanism, the fused query vector is updated using self-attention. The corresponding calculation expression is as follows: In the formula, Let be the updated query vector at time n. For layer normalization, It is a multilayer perceptron, and Self-Attenuation is a self-attention computation. This is the query vector at time n after fusion; Using a cross-attention mechanism, the updated query vector interacts with the fused features to dynamically update the query vector. The corresponding calculation expression is: In the formula, This is the updated query vector at time n+1. To fuse features, Cross-Atten extracts scene information relevant to the current decision by calculating the attention weights of the query and BEV features; The query vector is iteratively updated to obtain the final query vector, and then connected to the prediction head to predict multiple future waypoints and control commands.
7. An end-to-end autonomous driving system based on multimodal fusion according to claim 1, characterized in that, The output fusion module includes a trajectory prediction branch and a multi-step control prediction branch, and adopts a context-based fusion strategy to determine the final prediction result. The trajectory prediction branch uses a diffusion model to generate future trajectories, and then uses an MPC controller for lateral control and a PID controller for longitudinal control. The multi-step control prediction branch adopts a trajectory-guided attention mechanism, obtains information from the trajectory prediction branch, predicts the control actions of the next multiple steps, and directly predicts the vehicle's steering value, throttle value, and braking value. The context-based fusion strategy fuses the outputs of the trajectory prediction branch and the multi-step control prediction branch by weighted averaging, and adjusts the fusion weights according to the vehicle's driving state to obtain the final control signal.
8. An end-to-end autonomous driving system based on multimodal fusion according to claim 7, characterized in that, The multi-step control prediction branch uses a GRU model to output the updated hidden state based on the current features and the predicted control action. Hidden states of prediction branches using multi-step control Hidden states of trajectory prediction branches The attention map is calculated using MLP, and then... Image features F extracted by the attention map aggregation image feature extraction module MV Hidden states of multi-step control prediction branches The resulting concatenated feature is an enhanced feature used for predicting future control actions in the multi-step control prediction branch. The calculation expression for the enhanced feature is as follows: In the formula, To enhance the features, t represents time.
9. An end-to-end autonomous driving system based on multimodal fusion according to claim 7, characterized in that, The calculation rule for obtaining the final control signal based on the context fusion strategy is as follows: If the target vehicle is in a trajectory-specific scenario, then the expression for the final control signal α is: a=α·a ctl +(1-a)·a traj In the formula, α is the fusion weight, a ctl To control the output action of the prediction branch in a multi-step process, a traj The output action for the trajectory prediction branch; If the target vehicle is in a dedicated control scenario, then the expression for the final control signal α is: a=α·a traj +(1-a)·a ctl 。 10. An end-to-end autonomous driving system based on multimodal fusion according to claim 1, characterized in that, The system also includes a central hardware trigger, a time stamp module, a data buffer module, and a data alignment module; The central hardware trigger is used to send synchronization signals to each sensor in the environmental perception module, thereby triggering the operation of each sensor in a unified manner. The time stamping module is connected to the environment sensing module and is used to mark the environmental information data of each modality with timestamps. The data buffer module is connected to the time stamp module and is used to store environmental information data with timestamps. The data alignment module is connected to the data buffer module and is used to find and align the environmental information data of each modality in the data buffer module according to the timestamp.
Citation Information
Cited By
Intelligent driving end-to-end decision model iteration method and device, equipment and medium
CN121596752A
Displacement monitoring compensation method, training method and device based on image spatial-temporal characteristics
CN121616597A
End-to-end automatic navigation method, system, device, medium and program product
CN121804496A